More than 50 platforms compete for a listing in Gartner Peer Insights’ Audit Management Solutions market, and Gartner’s own 16 September 2026 release, drawn from 259 audit departments, found that 84 percent of audit functions have already adopted some form of audit management software, then made the sharper point: “it’s not a question of whether audit teams need” the software, but whether they capture the value they expected from buying it. A demo is usually where that gap opens. Left to run its own script, a vendor’s sales engineer shows the audit universe screen that looks best in a fresh browser tab, skips the SOX deficiency workflow that takes four clicks to reach, and answers an AI question with a feature name instead of a straight statement of which model runs it.
None of that is dishonest, exactly — it is what a sales demo is built to do. The fix is not a better vendor. It is a script the buyer writes once, runs identically with every finalist, and scores the same way every time. This guide is that script: 25 scenarios spanning planning, fieldwork, workpapers and review, issues, reporting, SOX, analytics, AI, administration and the auditee portal, each naming what to ask a vendor to show live, what a capable product looks like doing it, which of this site’s 12 scorecard areas it tests, and the red flag that means a vendor is narrating rather than demonstrating. It is the second stage of the site’s vendor-neutral RFP method, built to run once the RFP has narrowed a long list to two or three finalists, and it follows the same scorecard and fit-by-situation method used across every review and comparison in the independent buyer’s guide.
How to read this guide. Twenty-five scripted demo scenarios, organized by audit stage, built to run identically with every finalist vendor so the comparison is between products, not between sales teams.
What this guide is for. Running the live-demo stage of an audit software purchase, after the RFP has narrowed the field, with one script every vendor answers the same way and one scorecard that turns 25 live answers into a comparable number.
Evidence. Research-based: vendor documentation and release notes, public procurement records, third-party pricing data, verified user reviews on Gartner Peer Insights and G2, and analyst coverage, used to build scenarios that target real, documented gaps between vendor claims and what a product can actually show. We have not used any of these products hands-on for this guide.
Last verified. 27 September 2026.
In this guide
- Why a scripted demo beats a vendor-led pitch
- Demo logistics: your data, two hours, one script
- Planning and risk assessment scenarios
- Engagement workflow, workpapers and review scenarios
- Issues, reporting and SOX scenarios
- Analytics and AI scenarios
- Auditee portal and administration scenarios
- The 25-scenario demo scorecard
- Worked example: Lakeshore Bancorp runs three demos on one script
- Questions about the demo script
- Sources and verification
- Related guides
Why a scripted demo beats a vendor-led pitch
A vendor-led demo is built for one outcome: a good next meeting. That is a reasonable goal for a sales team and a poor one for a buyer, because it rewards whichever screen looks best in ninety seconds over whichever workflow will actually carry hundreds of engagements a year. A sales engineer who steers toward the strongest module is not being dishonest — that is the job — which is exactly why the buyer, not the vendor, has to bring the structure the session would otherwise lack.
A script fixes that without turning the session adversarial. Every vendor answers the same 25 prompts, in the buyer’s own data, in the same two-hour window, scored on the same sheet. What comes out the other side is not “which vendor gave the better pitch” but a set of area-by-area scores, the same 12 areas the site’s review method uses across every product it covers, that a CAE can defend to a CFO or an audit committee without relying on a feeling from the room.
Demo logistics: your data, two hours, one script
Five decisions determine whether a demo actually tests anything, and all five get made before the first vendor logs in.
- Your data, not the vendor’s sandbox. Send a de-identified extract of your own audit universe, a real risk and control matrix, and one real or realistic issue ahead of the session, and require the vendor to load it live. A vendor’s canned demo data is built to look good; yours is not.
- Two hours, not ninety minutes. Twenty-five scenarios do not fit into ninety minutes without skipping half of them, and a half-day session fatigues the panel more than it teaches them. Two hours, with a ten-minute buffer for follow-up questions, is enough for a vendor that actually built the product to show it.
- The same script for every vendor, the same week if the calendar allows it. A comparison made a month apart, after two vendors have started to blur together, is a comparison made from memory rather than from notes.
- Who attends. A CAE or audit lead to keep score, one or two staff auditors who will use the product day to day, and someone from IT or security for the administration scenarios; a SOX-heavy or regulated function should add its SOX lead for the SOX scenarios specifically.
- Record every session. A recording settles disagreements about what a vendor actually showed versus what one panelist remembers, and it lets anyone who missed the live session score from the same evidence as everyone else.
One scorecard area deliberately has no scenario below: area 12, cost and contract. A demo can perform a feature; it cannot perform a real price. Pricing and negotiation belong to the RFP stage and the conversation that follows it, covered in the site’s guides to audit software pricing and audit software due diligence, not to a two-hour session the vendor controls the room for.
Planning and risk assessment scenarios
Planning is where a demo script earns its keep fastest, because a canned risk model is one of the easiest things to fake with a slide and one of the hardest to fake live. The three scenarios below ask a vendor to build a real, editable risk-based plan from the buyer’s own audit universe, not to narrate one it prepared earlier.
| Scenario | Ask the vendor to show | What good looks like | Scores | Red flag |
|---|---|---|---|---|
| 1. Risk-based annual plan, built live | Load a sample of 20 to 30 entities from your own audit universe and generate a risk-scored annual plan in the session, not a plan prepared beforehand | Risk weights are visible and editable, and the ranked plan re-sorts instantly when a weight changes | Area 1 | The vendor switches to another client’s pre-built plan, or the risk score is a black box no one in the room can explain |
| 2. Resourcing against the plan | A resource-loaded view: hours available per auditor set against the plan’s estimated hours, and what happens when one engagement slips | A visual capacity view, not a spreadsheet export, that updates when dates move and shows utilization by auditor | Area 1 | “We would need professional services to build that view”, or resourcing lives only in a static report |
| 3. Mid-year universe change | Add one new auditable entity — a new product line or an acquired subsidiary — and show it flow into the risk assessment and the next plan revision | The new entity inherits the same risk taxonomy, generates a re-ranked plan section, and the change is timestamped | Area 1 | Adding an entity requires a support ticket or a data reload the vendor cannot do live |
The second and third scenarios matter more than they look. A vendor that handles resourcing gracefully has usually thought about the function as an operating team, not just a document repository; a vendor that stumbles on a mid-year universe change is telling the room something about how much of its “flexible” risk model is configuration versus custom development billed later. The site’s guide to the annual risk assessment covers the methodology a vendor’s plan-building screen should support without a workaround.
Engagement workflow, workpapers and review scenarios
This is the longest block of scenarios because it is the part of the product an audit team lives in every day: the planning memo, the risk and control matrix, test steps, evidence, review notes and sign-off. A vendor that is strong here and weak everywhere else is still a defensible choice for many functions; a vendor that is weak here rarely deserves the rest of its feature list.
| Scenario | Ask the vendor to show | What good looks like | Scores | Red flag |
|---|---|---|---|---|
| 4. Memo to RCM handoff | Start from a planning memo template and show its objectives and scope flowing into the RCM for the same engagement, live | Memo fields map directly into RCM rows without re-typing | Area 2 | The RCM is “a separate Excel upload” with no structural link to the memo |
| 5. Test step execution and evidence capture | Walk one test step from assigned to complete: attach evidence, log an exception, send it to a reviewer’s queue | Evidence attaches inline, exceptions are flagged distinctly, and a reviewer sees the item without navigating away | Area 2 | Evidence lives behind a SharePoint link pasted into a text field |
| 6. Engagement status and timeline | A portfolio view of every open engagement against its planned dates, then a drill-down into one that is late | A live status board, not a manually updated field, with a drill-down that shows the actual bottleneck step | Area 2 | Status is a dropdown someone updates by hand with no timeline view behind it |
| 7. Review notes and sign-off chain | A workpaper moving through preparer, reviewer and second-reviewer sign-off, including one open review note and its resolution | Sign-off is a locked, timestamped workflow state; review notes are threaded and can be resolved but not deleted | Area 2 | “Sign-off” is a checkbox with no trail of who signed which version |
| 8. Workpaper version history and audit trail export | Open a workpaper edited three times and export its full version history and audit trail in one action, the kind of evidence engagement documentation under Standard 14.6 requires | Every edit, by whom and when, is retrievable and exportable without help from IT | Area 3 | Version history exists on screen but cannot be exported, or only an admin account can pull it |
| 9. QAIP conformance metrics | A department-wide dashboard of quality metrics — on-time completion, review-note aging, methodology conformance — native to the platform, the kind of evidence Standard 12, Enhance Quality, asks a function to maintain | Metrics derive from real workflow data and export cleanly for a board packet or an external quality assessment | Area 9 | “QAIP” gets a shrug, or a pointer to a separate module priced extra with nothing to show live |
Scenarios 8 and 9 are the two most commonly skipped in a vendor-led pitch, because neither is a feature a sales engineer leads with unprompted. The site’s Global Internal Audit Standards reference map explains where Standard 14.6 and Standard 12 sit inside the broader framework, and its QAIP playbook is worth reading before the demo, not after, so the panel knows what a real conformance dashboard should be able to produce.
Issues, reporting and SOX scenarios
These seven scenarios cover what happens after fieldwork finds something: turning it into a tracked issue, following it to closure, reporting it upward, and, for a SOX-in-scope function, evaluating whether it rises to a deficiency. The follow-up scenarios track to the monitoring standard, 15.2, one of the areas examiners and audit committees scrutinize hardest.
| Scenario | Ask the vendor to show | What good looks like | Scores | Red flag |
|---|---|---|---|---|
| 10. Issue creation from a finding | Convert a fieldwork exception into a formal issue with a rated severity, a root cause and a target date, live | A direct link from the workpaper exception to the issue, with structured severity, not free text | Area 4 | Issues live in a separate module with no link back to the workpaper that raised them |
| 11. Validation, aging and escalation | An overdue issue: how aging is calculated, who gets escalated to, and what a validator sees before closing it | Automatic, department-wide aging; configurable escalation rules; validation that requires evidence, not a checkbox | Area 4 | Aging is a manual field, or escalation is “an email we would set up outside the system” |
| 12. Management self-reported status | A process owner updating remediation status on an open issue, and what the audit team sees on their side | A real self-service update path with a notification to the audit team, not a status the auditor has to chase and re-key | Areas 4 and 10 | All status changes route through the auditor; management has no direct path |
| 13. Engagement report generation | Generate a full report from workpaper and issue data, export it, then edit a finding’s wording without breaking the link to source data | The report pulls live from the engagement and re-exports cleanly after an edit | Area 5 | The report is “assembled outside the system” or breaks formatting on export |
| 14. Audit committee dashboard | A board-level dashboard: open issues by severity, plan status, one drill-down into a specific finding | A dashboard built for a non-auditor reader, refreshable live, screen-shareable without reformatting | Area 5 | “Dashboard” means exporting the data and building it yourself in another tool |
| 15. SOX control library and testing cycle | The control library for one process, a testing cycle assigned to it, and a sample pulled with the tool’s own sampling method | Key and non-key status, frequency and testing attributes live natively; sampling is documented and defensible | Area 6 | Sampling is “usually done in Excel” even inside a dedicated SOX module |
| 16. Deficiency evaluation and aggregation | A failed control test walked through deficiency evaluation to a materiality conclusion, with individually immaterial deficiencies aggregated | A structured evaluation workflow with a visible aggregation view across controls, not a memo written outside the system | Area 6 | Deficiency evaluation is a free-text field with no aggregation view behind it |
Scenario 12 is worth watching closely for any function that reports to a bank examiner or an external auditor: a platform with no real self-service update path for management usually means the audit team becomes the unpaid data-entry clerk for every status change, a pattern the site’s guide to issue log fields was built to catch on paper before it shows up live on a screen.
Analytics and AI scenarios
AI is where a scripted demo matters most and a vendor-led pitch is least trustworthy, because “AI” is the single easiest word in this market to say without meaning anything specific behind it. Gartner’s Market Guide for Audit Management Software, published 13 April 2026 by analyst James Bourke, tells buyers directly to scrutinize vendor claims about AI capabilities and be “particularly wary of agent-washing” — a label for automation dressed up as AI, or a demo slide standing in for a feature that does not yet exist in production. The three AI scenarios below answer that warning with a specific ask, not a general one.
| Scenario | Ask the vendor to show | What good looks like | Scores | Red flag |
|---|---|---|---|---|
| 17. Connect a source, run a scripted test | Connect a sample data extract, live, and run one scripted analytic test end to end in the session | The connection and the script run live, with a visible, exportable exception population | Area 7 | “Analytics” turns out to mean an integration with a separate tool the vendor does not own or demo |
| 18. Continuous monitoring and alerting | A scheduled test firing on a recurring basis and an exception alert reaching a named recipient | A real scheduling and alerting mechanism inside the platform, with a visible run history | Area 7 | “Continuous monitoring” turns out to be a quarterly manual re-run |
| 19. AI-drafted output with a checkable citation | Ask the AI feature to draft something real — a test step, a risk description, a workpaper summary — from your own uploaded document, and show the source behind each sentence | Every AI-drafted sentence traces to a specific, checkable source, clearly labeled as unreviewed until a human accepts it | Area 8 | No visible sourcing, or the sales engineer cannot say which model produced the answer |
| 20. Model name and data-use answer | Ask directly: which model runs this feature, is customer data used to train it, and what is the retention period | A named model and a written data-use statement, on the spot, not a slide read aloud | Area 8 | The sales engineer cannot name the model, or promises to “follow up” on a training-data question |
| 21. An agent that acts, and the gate in front of it | Any feature described as “agentic” actually taking an action — creating an issue, sending a notification, updating a field — with the human approval step shown before it commits | A visible, mandatory approval step before the agent touches real data, and a log of what it did and who approved it | Area 8 | The agent “just does it” with no approval gate, or the demo shows a slide instead of the agent running |
Scenario 20 has a real, sourced answer to compare against. Diligent’s own release notes for ACL AI Studio state plainly that it “now uses the Claude 4.6 model” — a specific, checkable answer, not a reference to unnamed “proprietary AI.” Optro (formerly AuditBoard) is a useful test case for scenario 25 below: this site’s own review work found its trust pages state that hosting “meets FedRAMP moderate impact requirements” without the platform actually holding a FedRAMP authorization — softer language than it first sounds, and exactly the kind of gap a direct question catches live that a slide would not. The site’s own guide to evaluating AI in audit software goes deeper on what “real” looks like in 2026 across every platform this site tracks.
Auditee portal and administration scenarios
The last four scenarios cover the two areas a vendor-led pitch skips most reliably because neither is glamorous: what the audit function’s own auditees experience day to day, and what IT has to sign off on before anyone else touches the product.
| Scenario | Ask the vendor to show | What good looks like | Scores | Red flag |
|---|---|---|---|---|
| 22. Auditee request and response portal | Log in as, or simulate, an auditee receiving a document request, uploading a response and marking it complete | A genuinely simple auditee-facing view, not the full auditor interface with some buttons hidden, with status visible to both sides | Area 10 | Auditees are expected to email documents that someone on the audit team re-uploads by hand |
| 23. Auditee action-plan updates | A process owner updating an action plan’s target date or status, and the notification the auditor receives | An in-system update with an automatic notification and a visible history of date changes | Area 10 | Date changes are silent and untracked, or require the auditor to notice on their own |
| 24. SSO, API and one live integration | Single sign-on working against a test identity provider, not just a settings screen, and one real API call, such as pulling a list of open issues | SSO authenticates live; the API returns real data in a visible response, with documentation the buyer can take away | Area 11 | “SSO is supported” with no live test, or API access sits behind an undisclosed higher tier |
| 25. Export and the vendor’s own security posture | A full engagement export, including workpapers and attachments, in a usable format, and the vendor’s current SOC 2 report or FedRAMP/GovRAMP status in writing | A real, complete export the buyer could open in another tool, and a specific, current certification answer | Area 11 | Export is “available on request” rather than demonstrated, or the security answer is a marketing claim rather than a document |
Scenario 22 is worth running even for a buyer who thinks of the demo as an audit-team decision alone. The PBC request list survival guide exists because the auditee side of a bad workflow is where an audit function’s internal goodwill gets spent, one unanswered reminder email at a time; a portal that fails scenario 22 in the demo will fail it in production too.
The 25-scenario demo scorecard
A script only helps if the scoring is consistent from vendor to vendor and from panelist to panelist. Copy the sheet below for each finalist and fill it in during or immediately after the session. The requirements matrix workbook carries the same scorecard as a ready-made sheet, with the weights and area averages calculated for you.
The 25-scenario demo scorecard
Score each scenario 0 to 3. 0: not shown, or the vendor could not produce it live. 1: shown, but needed a workaround, the wrong screen, or looked unfinished. 2: shown cleanly and matches what the RFP or the vendor’s own documentation claimed. 3: shown cleanly and exceeds the ask, with a specific, named example from the vendor’s own client base.
Total by scorecard area, not by vendor charisma. Add the scenario scores inside each of this site’s 12 scorecard areas — the sections above map directly, three planning scenarios into area 1, six fieldwork and review scenarios into areas 2, 3 and 9, and so on — and translate each area’s total into Strong, Adequate, Limited or Not offered the same way this site’s own reviews do.
Weight the areas before the first demo, not after. A SOX-heavy public company should weight area 6 double before comparing totals; a small team buying a first system should weight areas 10 and 11 higher and area 7 lower. Decide the weights on paper before anyone has watched a single vendor perform, so the vendor who happened to present best does not quietly move the weights afterward.
Carry two numbers forward, not one. The weighted total, and the count of scenarios scored zero. A vendor can post a respectable weighted average while quietly skipping four or five scenarios outright; that count belongs in the next round of RFP follow-up questions, not lost inside an average.
One shared document, one column per vendor. Score live or immediately after each session, not reconstructed from memory a week later once the finalists have started to blur together.
Worked example: Lakeshore Bancorp runs three demos on one script
Lakeshore Bancorp, the site’s recurring fictional $9 billion regional bank with 60 branches and 212 key controls, is replacing a spreadsheet-based SOX and audit process and has used the site’s RFP method to narrow a dozen vendors to three finalists. Its four-person panel — the CAE, the SOX lead, one staff auditor and a representative from IT security — books three two-hour sessions in the same week, each running the identical 25-scenario script against an extract of Lakeshore’s own control population and a sample audit universe of 40 branches and functions.
Because Lakeshore is a bank, the panel weights areas 6, 11 and 5 above the default before the first session, per the scorecard’s own instruction to weight first and watch second. Scenario 25 matters more here than it would for a smaller, unregulated buyer: the panel already knows, from this site’s own review work, that platforms including Onspring and TeamMate hold an actual FedRAMP Moderate authorization listed on the FedRAMP Marketplace, and that Diligent One documents FedRAMP Moderate and DoD IL5 configurations for its GovCloud region, while at least one other major platform’s trust pages use softer “meets FedRAMP moderate impact requirements” language without holding one — so scenario 25 is not a fishing expedition, it is a chance to watch each finalist confirm or contradict what its own public documentation already says.
All three finalists score respectably, and none sweeps every area, which the panel treats as normal rather than as a flaw in the script. Finalist A posts the strongest planning and SOX totals but takes two flat zeros in the AI table, once when its sales engineer cannot name the model behind a “smart insights” feature and once when an agentic action has no visible approval gate; the scorecard’s own rule — carry the zero count forward, not just the average — keeps those two zeros from disappearing into an otherwise strong score. Finalist B is weaker on deficiency aggregation but produces a complete, live engagement export in scenario 25 and a specific, current security answer on the spot. Finalist C has the softer FedRAMP language, loses points there, and wins the reporting scenarios outright with a board-ready dashboard the CAE could screen-share to Lakeshore’s audit committee unedited.
Lakeshore’s decision folds the weighted scorecard, the RFP method’s reference checks, and Finalist B’s live export into the same board-ready case the site’s guide to the business case for audit software walks through, and its first 120 days after signature follow the site’s guide to implementing audit management software. A demo score is an input to that decision, not the decision itself.
Questions about the demo script
How long should a vendor demo actually take?
Two hours per vendor is enough to run all 25 scenarios without rushing. Sessions under 90 minutes usually mean scenarios get quietly skipped, and sessions past three hours fatigue the panel more than they add useful evidence.
Should the vendor see the script in advance?
Yes. Send the script and the sample data ahead of time so the vendor’s own team can prepare a genuine live walkthrough instead of failing cold. The discipline is in running the identical script for every finalist and scoring what actually happens on screen, not in surprising anyone.
What if a vendor refuses to use our data?
Treat the refusal as data in itself. A vendor unwilling to load a sample of a buyer’s own audit universe, risk and control matrix or issue log into a two-hour session is telling the buyer something about how implementation will go once the contract is signed.
Do we need a separate demo just for AI features?
No. The three AI scenarios sit inside this same script specifically so AI claims get tested against the same one-ask, one-model, one-approval-gate standard as everything else, in the spirit of Gartner’s own warning to be wary of agent-washing rather than as a separate marketing conversation held on the vendor’s terms.
How is this different from the RFP?
The RFP narrows a long list to two or three finalists on paper. The demo script tests whether what those finalists wrote in the RFP is actually true on a screen. Skipping straight to a demo without an RFP first usually means comparing polish instead of substance.
What happens after the demo scores are in?
Reference checks, a closer security and AI data-use review, and a business case for the CFO or the board. The site’s guides to audit software due diligence, implementation and the business case each cover one of those next steps in full.
internalauditguide.com has no commercial relationship with any vendor named on this page. We take no vendor money, run no affiliate links and accept no sponsored placements, and no vendor saw this page before publication. Product and company names are the trademarks of their owners. Corrections: desk@internalauditguide.com.
Sources and verification
- Gartner press release, 16 September 2026 — the 84 percent adoption figure and the satisfaction-gap statement (accessed 27 September 2026).
- Wolters Kluwer, citing Gartner’s Market Guide for Audit Management Software (13 April 2026, James Bourke) — the agent-washing warning (accessed 27 September 2026).
- Gartner Peer Insights: Audit Management Solutions market — the scale of the market this script is built for (accessed 27 September 2026).
- FedRAMP Marketplace: Onspring GovCloud — a real FedRAMP Moderate authorization, used as the worked example’s contrast (accessed 27 September 2026).
- Diligent One Platform Trust and Compliance — FedRAMP Moderate and DoD IL5 authorization detail (accessed 27 September 2026).
- Diligent help center: ACL AI Studio release notes — the named-model statement used in scenario 20 (accessed 27 September 2026).
- Optro: AI and LLM information page — the FedRAMP language example used in the AI scenarios and the worked example (accessed 27 September 2026).
Related guides
- Internal audit software: the independent buyer’s guide — every review, comparison and buying guide in one place.
- How we review audit software — the evidence levels, the scorecard and the fit-by-situation method.
- The audit software shortlist finder — eight questions, a shortlist with the reasons from each review.
- The requirements matrix — 156 weighted requirements and vendor scoring in a free Excel workbook.
- Selecting an audit management system — the vendor-neutral RFP method this script picks up after.
- Types of internal audit software — audit management, GRC, SOX, compliance and analytics, defined.
- Best internal audit software — 25 platforms and tools compared by use case.
- Internal audit software pricing — real numbers, pricing models and how to negotiate.
- Audit software due diligence — security, data residency, AI data use and vendor stability.
- Evaluating AI in audit software — what is real and what is agent-washing.
- The business case for audit software — getting budget approved by the CFO and the board.
- Implementing audit management software — the first 120 days and migrating off Excel.
- 15 mistakes internal audit teams make when buying software — and how to avoid each one.
- Audit software for small internal audit teams — options and real prices for one to five auditors.
- Audit software vs Excel and SharePoint — when to switch, and how big you need to be.
- GRC suite vs standalone audit management software — how to make that call before shortlisting anything.
- Internal audit software vs compliance automation — why Vanta and Drata are not on this shortlist.
Leave a Reply