Strip an audit to its skeleton and only one thing is load-bearing: the evidence. The plan, the walkthroughs, the workpapers, the meetings — all of it exists to produce support for conclusions, and a conclusion is worth exactly what supports it. Yet evidence is the part of audit training that gets a slide, not a discipline: new auditors learn to fill templates long before they learn to ask whether the screenshot they just pasted proves anything at all. The result shows up at review — files heavy with exhibits and light on proof, conclusions that outrun their support, and the reviewer’s least favorite discovery: an opinion resting entirely on what management said.
This guide is the discipline. What sufficiency and appropriateness actually mean, the five types of evidence and what each can and cannot prove, the reliability hierarchy and the modifiers that move evidence up and down it, corroboration rules of thumb you can apply mid-fieldwork, one assertion tested three ways — weak, adequate, strong — with the conclusion each version can honestly support, and the review-note classics. The Global Internal Audit Standards set the bar in four words — information that is sufficient, reliable, relevant, and useful — and everything below is those four words made operational.
This guide was rewritten in September 2026 to add what the August version left out: the evidence each assertion needs, an evidence plan template for writing procedures that name their proof in advance, the evidence behind eight conclusions from the engagements on this site with the rung each one stood on, the problems that screenshots, exports and generated documents have introduced since the hierarchy was first written, the evidence standards for issue validation and referrals, and the links to the guides that apply all of it. The definitions, the five types, the hierarchy, the corroboration rules and the three-way example are unchanged. Standard 14.1 of the Global Internal Audit Standards requires information that is sufficient, reliable, relevant and useful, and this guide remains those four words made operational.
In this guide
- Sufficiency and appropriateness, precisely
- The five types of evidence — and what each proves
- Evidence by assertion: what each claim needs
- The reliability hierarchy
- How much is enough: sufficiency by risk and by type
- Corroboration rules of thumb — and the direction of testing
- The evidence plan: naming the proof before the fieldwork
- One assertion, three ways: weak, adequate, strong
- Worked example: the evidence behind eight conclusions
- Screenshots, exports and generated documents: the hierarchy in 2026
- Evidence standards for validation and for referrals
- The review-note classics
- Where to go next
Sufficiency and appropriateness, precisely
Sufficiency is quantity: is there enough evidence to support the conclusion at the level of assurance you are claiming? Enough is a function of two things — the risk riding on the conclusion (a finding that will trigger a restatement conversation needs more than a routine control pass), and the quality of each piece, because better evidence means less of it is needed. That trade is the engine of efficient fieldwork: one reperformed reconciliation outweighs five interviews about reconciliations, and the auditor who understands the exchange rate stops overcollecting weak evidence to compensate for its weakness — a stack of low-grade exhibits does not compound into proof. Appropriateness is quality, and it splits into two tests that must both pass. Relevance: does this evidence bear on the specific assertion being tested — not the process in general, this claim? Evidence that a control exists is not evidence it operated; evidence about March is not evidence about the year. Reliability: can the evidence be trusted — who created it, who has been able to alter it, and how did it reach you? Relevance without reliability is a rumor on point; reliability without relevance is a certified answer to a different question. Every review note about evidence is, underneath, one of these three words — not enough, not on point, or not trustworthy.
The five types of evidence — and what each proves
| Type | What it is | What it proves well | Its structural weakness |
|---|---|---|---|
| Inquiry | Asking people — interviews, walkthrough conversations, questionnaires | Direction: where to look, how the process claims to work, what changed | It is testimony. People are wrong, optimistic, or interested — the same epistemology problem as the RCSA. Inquiry alone never supports a conclusion |
| Observation | Watching the process or control happen | That it happened then, performed that way while watched | Proves the moment, not the period — and observed behavior is best behavior |
| Inspection | Examining documents, records, system configurations, or physical assets | Existence and terms of the thing inspected; configuration states; approvals present | Reliability inherits from provenance — who made the document, and could it have been altered on its way to you? |
| External confirmation | Evidence obtained directly from an independent third party | Balances, terms, and facts held by outsiders with no stake in your conclusion | Scope-limited and slow; confirming parties answer exactly what was asked, nothing more |
| Reperformance / recalculation | Independently redoing the control or computation and comparing outcomes | The strongest form: you generated the evidence yourself, so provenance is not in question | Expensive per unit; only as meaningful as the inputs you reperformed from |
| Analytics | Computations across a full population — patterns, outliers, reconciliations at scale | Population-level claims no sample can make; the anomalies worth the file review | Worth nothing more than the data’s integrity — an unverified extract analyzed brilliantly is brilliant nonsense |
The craft is matching type to claim. Existence questions want inspection and confirmation; process-understanding questions want inquiry and observation; operating-effectiveness questions want reperformance and inspection of the control’s own residue — the attribute logic from our TOD versus TOE guide; population claims want analytics with the data’s integrity established first. Most weak files are not short of evidence — they are full of the wrong type for the assertion at hand.
Evidence by assertion: what each claim needs
Most weak files are full of evidence for a different assertion than the one being tested, and the cure is to name the assertion before choosing the evidence. The table gives the claims auditors most often make, the evidence type and direction each needs, and the evidence that is routinely filed for it and does not work. The financial statement assertions are included because internal auditors test them in every process audit whether or not they use the vocabulary; the control assertions are the ones the TOD versus TOE guide separates.
| Assertion | What it claims | Evidence that supports it | Evidence filed for it that does not |
|---|---|---|---|
| Existence and occurrence | Recorded items are real and happened | Vouch from the record to the source document, the counterparty or the physical asset; confirm externally where the counterparty is independent | A system report listing the items; management’s assurance that they exist |
| Completeness | Nothing that should be recorded is missing | Trace from an independent source population (shipments, contracts, terminations, bank statements) into the records; reconcile control totals | Vouching from the records, which cannot find what is not there |
| Accuracy and valuation | Amounts and estimates are right | Recalculate from source inputs; re-perform the estimate from data; compare to external prices or subsequent outcomes | The reviewer’s sign-off on the calculation; the prior year’s figure |
| Cut-off | Items are in the right period | Two-sided tests around the period end from independent date evidence: delivery records, bank dates, carrier data | The posting date, which is what is being tested |
| Rights and obligations | The organization owns the asset or owes the liability | Contracts, title documents, confirmations from the counterparty or custodian | The asset’s presence on the register |
| Control design | The control, if it operates, would address the risk | Walkthrough of one instance with the configuration, the procedure and the evidence it leaves; a test of one | The policy document alone |
| Control operation | The control operated as designed throughout the period | Inspection of the control’s own residue (approvals, reconciliations, review evidence) for a sample across the period, or re-performance, or full-population analytics on system logs | Inquiry of the performer; the control’s existence; one instance observed on the day |
| Segregation | Incompatible duties are held by different people | The system’s access and workflow configuration extracted and joined to activity logs | The organization chart; job descriptions |
| Population integrity | The data analyzed is the whole population | Reconciliation of the extract to an independent control total (ledger, bank, provider settlement) with the query and filters recorded | The extract’s row count; the auditee’s statement that it is complete |
Two disciplines follow from the table. The assertion is written at the top of the workpaper before the evidence is gathered, as the model workpaper shows, so that the evidence is chosen for it rather than fitted to it afterward. And a conclusion is written only for the assertion the evidence addresses: a file that tested existence has earned a sentence about existence, and the sentence about completeness waits for the trace.
The reliability hierarchy
Rank evidence by how hard it would be for the story to be wrong, and a ladder emerges that holds across almost every engagement:
- Auditor-generated — reperformance, recalculation, your own analytics on verified data. You made it; nobody could have staged it.
- External, obtained directly — confirmations and reports that travel from the third party to you without passing through the auditee’s hands.
- External, held by the entity — vendor invoices, bank statements, contracts: independent origin, but the copy you see spent its life in the auditee’s custody.
- Internal, system-generated under tested controls — reports and logs from systems whose change management and access you have assessed; their reliability is exactly as strong as that ITGC layer, the dependency our ITGC guide unpacks.
- Internal, manually produced — spreadsheets, memos, tracking logs: useful, alterable, and produced by people with a stake in what they show.
- Inquiry and representations — the floor of the ladder. Indispensable for direction, insufficient for conclusions.
Four modifiers move any piece up or down: originals beat copies (every generation of copying — the scan of the printout of the export — adds an opportunity for alteration and subtracts context); contemporaneous beats reconstructed (a log written at the time outweighs a summary assembled for the auditors); controlled systems beat spreadsheets (an approval in the workflow tool outranks the same approval in an emailed Excel); and disinterested beats interested (the same fact from someone with nothing riding on it is worth more). None of this makes low-rung evidence worthless — it sets the price: the lower the rung, the more corroboration the conclusion needs before it can bear weight.
How much is enough: sufficiency by risk and by type
Sufficiency is the question reviewers ask most and the guide answers least precisely, because the honest answer is a function rather than a number. The function has three inputs. The risk riding on the conclusion: a control the external auditor will rely on, a finding that will be rated High, a matter that may become a referral, each needs more than a routine control pass, and the sample size guide gives the arithmetic behind 25, 40 and 60 for the routine case. The quality of each piece: auditor-generated evidence on a full population needs no sample at all, because the question of how many was answered by “all”, while inquiry needs corroboration however much of it there is, because the type does not compound. And the assertion: completeness claims need a source population traced in full or a sample from it, existence claims can be sampled from the records, and operating effectiveness claims need the period covered, so that a sample of twenty-five drawn from one quarter supports a conclusion about that quarter. The table gives the combinations that recur, and the reviewer’s question for each is the same: does the evidence entitle the sentence written above it, at the level of assurance the report claims.
| Conclusion sought | Sufficient evidence | Not sufficient |
|---|---|---|
| A routine control operated throughout the year | 25 instances across the period from the sampling standard, inspection of the control’s residue, exceptions dispositioned | Five instances from the last month; observation of the control on the day of the walkthrough |
| A control the external auditor will rely on operated | The sample the reliance agreement specifies, or full-population analytics on the control’s log, with population integrity proved | The internal auditor’s judgment that it “seems to work” |
| A population statement (all changes, every payment) | The full population extracted, reconciled to a control total, analyzed by the auditor | An extract the auditee supplied, unreconciled, however large |
| An estimate is reasonable | Re-performance from source inputs and a back-test against outcomes over several periods | The reviewer’s sign-off and the prior year’s method |
| A High finding’s condition | Full population where the data allows, otherwise a sample sized for the risk, with every exception corroborated and management’s confirmation of the facts recorded | The exceptions in a sample of 25, unconfirmed, with a projection the sample cannot support |
| A matter for referral exists | The pattern documented from the records as found, preserved, with no further evidence gathered by the auditor | An interview with the person the pattern points to |
Corroboration rules of thumb — and the direction of testing
- Inquiry never concludes alone. Treat every interview answer as a hypothesis that names its own test — “we reconcile that monthly” is an invitation to inspect twelve reconciliations, not a fact for the file.
- Stack types, not instances. Three people agreeing is one type of evidence three times; an interview plus the system log plus a reperformance is corroboration. Independence between sources is what multiplies strength.
- Price the evidence to the fight. The more a finding will be contested — money, blame, a serious rating per the 5 C’s — the higher up the hierarchy its support must come from. Write the contested finding from reperformance and system records, never from recollections.
- Extracts need provenance. Any data-based conclusion inherits the integrity of the extract: who pulled it, from where, with what filters, reconciled to what control total. Population integrity is procedure one of any analytics work — the discipline our IAM guide applies before believing a single account listing.
- Screenshots need context. System, screen, date, account, and who captured it — or it is an anonymous picture of software.
And one rule that is really about relevance: the direction of testing must match the assertion. Selecting recorded items and vouching them back to source documents tests existence — everything recorded is real. It cannot test completeness, because items missing from the records were never in your selection pool. To test completeness you trace the other way: start from the source population — shipments, contracts, terminations — and follow it into the records, hunting for what never arrived. Files that vouch when they should trace produce the classic false comfort: a beautifully supported conclusion that the wrong question was answered thoroughly.
The evidence plan: naming the proof before the fieldwork
The work program guide requires each procedure to name its evidence in advance, and the evidence plan below is the form that requirement takes for a single test. It is completed at planning, reviewed by the engagement lead, and it is the reason a reviewer at the end of fieldwork finds a file whose exhibits support its conclusions rather than one whose conclusions were fitted to its exhibits.
Evidence plan for a test. 1. Assertion: the claim this test will conclude on, in one sentence, with the population and period. 2. Evidence type and direction: which of the five types, and for completeness claims, the independent source the trace starts from. 3. Source and provenance: where the evidence comes from, who produces it, whether it passes through the auditee, and the rung on the hierarchy it will occupy. 4. Population integrity: for any data-based test, the control total the extract will be reconciled to and who will run the query. 5. Sufficiency: the sample size and selection method from the sampling standard, or the statement that the test runs at full population, and the risk level that justifies it. 6. Corroboration: what second type of evidence will be obtained where the primary evidence is inquiry, observation or an internal document. 7. The sentence: the conclusion the evidence will entitle the auditor to write if no exceptions are found, and the conclusion if exceptions are found, drafted in advance. 8. What would change the plan: the exception rate or the finding that would require the sample to expand, the type to change, or the test to stop and become a referral.
Item 7 is the one that changes how auditors work. Drafting the conclusion’s sentence before the evidence is gathered forces the question the guide is about, what am I entitled to conclude from what I will actually have, to be answered while the plan can still change, and it makes the file review a comparison of two sentences rather than a search through exhibits.
One assertion, three ways: weak, adequate, strong
The assertion: “Terminated employees’ system access is removed within five business days.” Watch what each evidence package can honestly support.
Weak. Interviewed the IAM manager, who described the termination workflow; inspected the policy requiring removal within 5 business days. — Two pieces, both real, neither touching the assertion: the policy proves the rule exists, the interview proves management believes the process works. Honest conclusion available: “a process exists and is documented.” Anything stronger — “access is removed timely” — is the interview wearing a conclusion’s clothes.
Adequate. Selected 25 terminations across the period per the sampling standard; for each, inspected the HR termination record and the deprovisioning ticket; computed elapsed days; 24 within SLA, one at 7 days, dispositioned. — Type matches assertion (inspection + recalculation of timing), coverage spans the period, exceptions handled. Honest conclusion: “for the sampled population, removal operated within SLA with one exception” — a defensible operating-effectiveness conclusion at sample-level assurance.
Strong. Obtained the full HR termination population and the complete account listing directly from the systems; verified the extract’s completeness against payroll’s final-pay records; joined terminations to accounts and computed removal lag for every leaver, contractors included; cross-checked residual active accounts against authentication logs for post-termination use. — Auditor-generated, full-population, completeness-verified in both directions. Honest conclusion: “for the entire population, 94% of access was removed within SLA; 11 accounts remained active beyond it, two with post-termination logins” — a statement precise enough to price the risk, name the consequence, and survive any challenge management can mount. Same assertion, three files, three very different sentences you are entitled to write — and that entitlement is the whole game.
Worked example: the evidence behind eight conclusions
The engagements written up across this site produced findings that survived closing meetings, and they survived because of what stood under them. The table takes eight conclusions from those engagements and shows the evidence each rested on, its rung on the hierarchy, and the sentence the evidence entitled the auditor to write.
| Conclusion | Evidence | Rung | The sentence it earned |
|---|---|---|---|
| MidState’s vendor bank changes were not independently verified (AP audit) | All 46 changes from the ERP change log joined to the payment file; the call-back records requested for each; the bank’s letter about the stopped attempt | Auditor-generated, full population; external obtained directly | “31 of 46 changes had no call-back record; 14 preceded a large payment within ten days; one was an attempt the bank stopped” |
| MidState’s ghost employee indicators reduced to nine (payroll analytics) | HR master and payroll register extracted with completeness proved against the pay run totals; 168 hits corroborated to timekeeping and bank records | Auditor-generated on verified extracts | “Of 168 flagged records, nine remained after corroboration: six shared bank accounts, three paid without timekeeping activity” |
| Brightwater’s sole-source justifications did not demonstrate uniqueness (procurement audit) | The 57 memos inspected; the market check re-performed by the auditor for 15, with the searches and supplier correspondence filed | Internal manual documents, corroborated by auditor re-performance | “On 15 re-performed market checks, plausible alternatives existed for six, found within twenty minutes each” |
| Pennine’s close ran on day eleven, not day eight (close audit) | The task tool’s completion timestamps for twelve months, and one close observed live | Internal system-generated under tested access, plus observation | “The close completed on working day eleven in ten of twelve months against a calendar of eight” |
| Lakeshore’s certifications were signed over open exceptions (reconciliation program audit) | The reconciliation tool’s status at the certification date joined to the signed certifications | Auditor-generated join on system records | “22 owners certified accounts that carried unexplained differences older than ninety days at the certification date” |
| Brightwater’s trade-promotion accrual was one-sided (revenue audit) | Eight quarters of accrual balances against the subsequent deductions, joined by program | Auditor-generated back-test on ledger and deductions data | “Under-estimated in six of eight quarters by 180,000 to 420,000 dollars, over-estimated in none” |
| Brightwater’s idle bottling line was an impairment indicator (fixed asset audit) | Operations interviews, corroborated by the plant’s production log and a site inspection | Inquiry, corroborated by internal system records and observation | “A line with a net book value of 1.4 million dollars has been idle for eleven months with no restart plan” |
| MidState’s payroll provider’s complementary controls were not operated (payroll audit) | The SOC 1 report’s user control list, obtained from the provider; the absence of any assignment, review record or reviewer confirmed by inquiry and inspection | External held by the entity; corroborated by inspection of the absence | “Two of the four complementary user controls the report assumes MidState operates were not being performed” |
Seven of the eight rest on auditor-generated or externally obtained evidence, and the eighth, the idle line, rests on inquiry that was corroborated the same day by a production log and a walk to the plant floor. None rests on what management said, and every sentence in the last column contains a number the file can produce. That is what the hierarchy looks like when it is used rather than taught.
Screenshots, exports and generated documents: the hierarchy in 2026
Three kinds of evidence have become more common and less reliable since the hierarchy was first written. Screenshots are the first. A picture of a screen proves that a screen looked that way to someone at some moment, and nothing else; without the system, the account, the date, the filter and the capturer, it is an anonymous picture of software, and with image editing now trivial it is weaker than that. The fix is provenance, captured at the moment: the system’s own export with its metadata, or a capture with the address bar, the user and the timestamp visible, captioned by the auditor who took it. Exports are the second. A spreadsheet that arrived by email from the auditee is internal manual evidence however system-generated its origin, because the auditee could have edited it on the way; the fix is to run the query yourself, or to watch it run, or to reconcile the extract to a control total you obtained independently, and the access review guide‘s population-integrity step is the template. Generated documents are the third and the newest. Receipt images, invoices, confirmations and even correspondence can now be produced by tools that leave no obvious trace, which is why the travel and expense guide calls receipt images the weakest evidence class an auditor holds, and why the corroboration rule has hardened: a document the auditee supplied is corroborated by a record the auditee did not produce, the card feed, the bank, the carrier, the counterparty, or it is treated as inquiry in paper form. The hierarchy has not changed; the floor has moved up, and the auditor’s habit of asking who made this and who could have changed it has become the whole discipline rather than one rung of it.
Evidence standards for validation and for referrals
Two moments in a finding’s life carry their own evidence standard. Issue validation is the first: the issue validation guide sets the evidence closure requires by rating, re-performance on post-remediation data for a High, inspection of the implemented control with a sample for a Medium, management’s evidence reviewed for a Low, and the reason the standard is graduated is the hierarchy. A High closed on management’s memo has been closed on inquiry, the floor of the ladder, for the finding that most needed the top of it. Referrals are the second: the moment a pattern becomes a possible fraud, the evidence rules change from the auditor’s to the investigator’s, and the first 48 hours protocol covers preservation, chain of custody and the interview the auditor must not conduct. The evidence the auditor has already gathered remains valid; what changes is that no more is gathered by the auditor, and what exists is preserved in the state it was found, because evidence that is sufficient for a finding may be inadmissible for anything else if it was handled as a finding’s evidence after the moment it became something more.
The review-note classics
| The pitfall | Why it fails | The fix |
|---|---|---|
| Evidence by screenshot | No system, date, account, or capturer — an unverifiable picture | Caption every capture with provenance, or export from the system with metadata intact |
| The unverified extract | Analytics on a population nobody reconciled prove nothing about the real population | Population-integrity step first: completeness verified against an independent control total |
| Inquiry stacked on inquiry | Three interviews are one weak type, thrice — agreement among the interested is not corroboration | Cross the type boundary: pair testimony with records or reperformance |
| Vouching for completeness | Direction mismatch — you cannot find missing items by examining recorded ones | Trace from the source population toward the records when the claim is “nothing is missing” |
| Evidence for the adjacent assertion | Proof the control exists filed as proof it operated; March filed as the year | Write the assertion at the top of the workpaper and test the evidence against it, per the workpaper standard |
| Overcollection | Two hundred exhibits, no argument — sufficiency measured by folder size | Each piece earns its place by supporting a stated conclusion; delete what supports nothing |
Where to go next
Evidence discipline is a habit of asking one question relentlessly: what am I entitled to conclude from what I actually have? The hierarchy prices each piece, corroboration says how pieces combine, direction says whether the test faces the assertion, the evidence plan names the proof before the fieldwork, and the three-way example shows how far the same claim can stretch or snap. Write procedures that name their evidence in the work program, record the selection in the sampling memo, document to the standard the model workpaper shows, and let the strength of the file decide the strength of the sentence, which is the discipline the conditions guide applies to the sentence itself. Auditors are in one business, the business of being believed, and evidence is the entire inventory.
Related guides
- Annotated workpaper examples — the assertion at the top and the evidence beneath it
- Writing the audit work program — procedures that name their evidence in advance
- The sampling memo template — sufficiency decisions, documented
- Audit sample sizes demystified — where 25, 40 and 60 come from
- Test of design vs operating effectiveness — the two control assertions and their evidence
- How to perform a user access review — population integrity before believing a listing
- IT general controls for non-IT auditors — why system reports are only as reliable as the ITGC layer
- Writing airtight conditions — the sentence the evidence entitles you to
- The 5 C’s of audit findings — pricing evidence to the fight
- Issue validation in internal audit — the evidence standard by rating
- When internal audit finds fraud: the first 48 hours — when the evidence rules change
- How to audit travel and expense — receipt images as the weakest class
- The RCSA process — the epistemology problem of self-assessment
- The life of a finding — corroboration as the finding’s third stage
- Fieldwork and testing guides and workpapers and documentation guides — the full collections
- Finding Severity Ratings: The Calibrated Matrix — anchors, the grid, twelve cases and the rating memo.
Leave a Reply