Sampling decisions are made in the first hour of a test and defended in the last week of an external quality assessment, usually by someone who was not there. How was the population defined, and how do we know it was complete? Why forty and not twenty-five? How were the items chosen, and could anyone reproduce it? What counted as an exception, and what happened when one appeared? A workpaper that answers those questions in a paragraph at the top is a workpaper that survives review; one that leaves the reviewer to reverse-engineer the decisions from the test results is the reason so many file reviews end with an observation about sampling.
The sampling memo is one page, written before any item is tested, that records every decision in a fixed order: objective, population and its completeness, method, size and its rationale, selection mechanics, what an exception is, how results will be evaluated, and what the conclusion may say. This guide gives the template in full, annotates each section with what to write and what goes wrong, sets out the size rationales for the common cases, explains how exceptions are evaluated, shows how the outputs of the free mySampler tool drop into the memo, and works a completed memo from MidState Beverage’s route cash audit. It is the documentation companion to the sample size guide and the judgmental sampling guide.
In this guide
- Why a memo, and why before testing
- The sampling memo template
- Section by section: what to write and what goes wrong
- Size rationales for the common cases
- Evaluating results: exceptions, expansion, and the conclusion
- The monetary-unit variant of the memo
- Filling the memo from mySampler
- Worked example: MidState’s settlement reconciliation memo, completed
- Where the memo sits in the file, and how a reviewer uses it
- Common mistakes
Why a memo, and why before testing
The Global Internal Audit Standards require engagement conclusions to rest on sufficient, reliable, relevant, and useful information and require the work to be documented so that a reviewer can understand what was done and why. Sampling is where those two requirements meet: the sufficiency of the evidence depends on how the sample was designed, and the design is invisible unless it is written down. The external audit standards, AU-C 530 and the PCAOB’s AS 2315, spell out the same elements, population, method, size, selection, evaluation, and reviewers trained on either set look for them in that order.
Writing the memo before testing matters for a reason that has nothing to do with paperwork. A size chosen before the results are known is a judgment; a size chosen afterward is a rationalization, and the two are indistinguishable on the page unless the memo carries a date and a reviewer’s initials from before fieldwork. The same is true of the definition of an exception: decided in advance, it is a standard; decided when the first questionable item appears, it is negotiable, and the auditee will negotiate it. The memo turns sampling from something the auditor did into something the auditor decided, which is what a reviewer is assessing.
| Reader | What they look for in the memo | What its absence costs |
|---|---|---|
| Engagement reviewer | That the design fits the objective and the size is justified before results | Review notes, rework, and a test that may have to be redone |
| Quality assessor | The Standards’ documentation and evidence requirements met on the page | A file-level observation that repeats across engagements |
| External auditor considering reliance | Elements comparable to AS 2315 or AU-C 530, so the work can be used | Re-performance of the test at the company’s expense |
| Regulator or examiner | Population completeness and independence of selection | A finding about the audit function |
| The auditor, a year later | What was done, to repeat or to defend it | Reconstructing decisions from memory, which fails |
The sampling memo template
Sampling memo. Engagement: ________ Test reference: ________ Prepared by: ________ Date: ________ Reviewed by: ________ Date: ________
1. Test objective and the control or assertion tested. [One sentence: what the test will conclude on. “Whether the daily settlement reconciliation was prepared and independently reviewed for each route settlement in the period.”] Type of conclusion sought: [operating effectiveness rate / misstatement amount / targeted risk / design].
2. Population. Definition: [the set of items the conclusion will cover, with the period, the units, and any exclusions]. Source: [system, report name, parameters, extraction date, who extracted it]. Count and value: [n items; $ total]. Completeness check: [what the population was reconciled to, the result, and where the reconciliation is documented]. Stratification, if any: [strata and their counts and values].
3. Sampling method. [Random / systematic with a random start / stratified random / monetary-unit / targeted judgmental / full population.] Reason for the method: [why this method suits the objective and the population].
4. Sample size and rationale. Size: [n, or per stratum]. Basis: [the convention, the statistical parameters (confidence, tolerable rate, expected rate), or the judgmental criteria and coverage]. Risk assessment behind the size: [the control’s importance, the risk rating of the area, prior results, reliance placed on other testing].
5. Selection mechanics. Tool and settings: [mySampler / Excel formula / other; the random seed or start point recorded]. Selection performed by: [auditor name, date]. Items selected: [reference to the listing]. Independence: [the auditee’s role, if any, was limited to retrieving documents for items the auditor named]. Replacement rule: [what happens if a selected item cannot be tested: replaced by the next random item, or treated as an exception].
6. Definition of an exception. [What counts as a failure of the attribute tested, decided now: “the reconciliation was not prepared; or was prepared but not reviewed by someone other than the preparer; or the reviewer handled the cash for that settlement; or the review is undated or more than two business days after the settlement.”] What does not count: [documentation shortfalls that do not indicate control failure, and how they will be recorded].
7. Evaluation approach. Tolerable exception rate or amount: [n]. Action on exceptions: [each exception investigated for cause before counting; the sample expanded by [n] if [condition]; the control concluded ineffective if [condition]]. Projection: [whether and how results will be projected; none for targeted selections].
8. Conclusion boundary. [What the results will support and what they will not: population conclusion, targeted-risk conclusion, or design only.]
9. Changes after this memo. [Any deviation from sections 2 to 8 during testing, with the reason, the date, and reviewer approval. Blank means none.]
Section by section: what to write and what goes wrong
| Section | What to write | What goes wrong |
|---|---|---|
| 1. Objective | One sentence naming the attribute or amount and the conclusion type; the conclusion type decides everything below it | “Test invoices” as an objective; a design objective tested with a sample built for a rate |
| 2. Population | The set the conclusion will cover, its source with parameters, count and value, and the reconciliation that shows it is complete | A report accepted as the population without reconciling it to anything; two acquired depots missing because they are not on the ERP |
| 3. Method | The method and the reason it fits: random for a rate, monetary-unit for value, targeted for concentrated risk, full for small populations | “Judgmental” with no criteria; random from a population where twenty items are 70 percent of the value |
| 4. Size | The number and its basis: the convention for the control’s frequency, the statistical parameters, or the judgmental coverage; the risk factors that pushed the size up or down | A size chosen after the first ten items looked clean; no link between the area’s risk rating and the size |
| 5. Selection mechanics | The tool, the seed or start, who selected, the listing reference, the auditee’s role, the replacement rule | Items the auditee supplied; no seed, so the selection cannot be reproduced; unavailable items silently replaced by convenient ones |
| 6. Exception definition | What fails the attribute, in the words a reviewer can apply; what does not count and how it is recorded | Decided when the first questionable item appears; “minor” exceptions waved through without a rule |
| 7. Evaluation | Tolerable rate, the investigate-before-counting rule, the expansion trigger, the projection method or its absence | Expanding the sample until the rate looks acceptable; projecting from a targeted selection |
| 8. Boundary | What the conclusion may say, written now | A population conclusion drawn from the ten largest items |
| 9. Changes | Every deviation, dated and approved | Silent changes to the population or the size mid-test |
The population section is where most memos fail and where the most effort belongs. A population is not a report; it is a report that has been reconciled to something independent of it, a control total, a ledger balance, a headcount, a prior-period count plus movements, and the reconciliation lives in the file. The IPE guide covers the completeness and accuracy testing that turns a report into a population, and the random sampling guide covers the tie-out mechanics in Excel and Sheets.
Size rationales for the common cases
Section 4 is a sentence or two, but the sentence has to come from somewhere, and the table gives the rationale for each common case in the form the memo needs. The conventions are those the sample size guide explains; they derive from attribute sampling at high confidence with a tolerable rate around ten percent and zero expected exceptions, which is why one exception changes everything.
| Case | Size | Rationale to write in section 4 | Adjust upward when |
|---|---|---|---|
| Annual control | 1 | “The control operates once; the single instance is tested in full with re-performance.” | Never; re-perform instead |
| Quarterly | 2, or all 4 for key controls | “Two of four instances per the function’s convention for quarterly controls; all four because the control is key to [assertion].” | Key control; prior exceptions |
| Monthly | 2 to 4 | “Three of twelve per the convention for monthly controls, increased from two because the area is rated high risk.” | High risk; new control; prior exceptions |
| Weekly | 5 to 10 | “Eight of 52 per the convention, the upper end because the control was new in the period.” | New control; turnover in the operator |
| Daily | 20 to 25, up to 40 | “Twenty-five of approximately 250 per the convention for daily controls; 40 where the control is the primary control over [risk].” | Primary control; high risk; no compensating control |
| Many times daily | 25 to 40, up to 60 | “Forty of 37,400 per the convention for high-frequency controls, stratified by depot so that every location is represented; 60 where the control is primary and the area high risk.” | Primary control; high risk; heterogeneous population |
| Monetary-unit (value conclusion) | Calculated | “Sample size of [n] from the MUS parameters: population value [$], tolerable misstatement [$], expected misstatement [$], confidence [percent]; interval [$].” | Lower tolerable misstatement; higher expected error |
| Targeted judgmental | By criteria | “All items meeting [criteria], [n] items covering [percent] of value; no population rate will be concluded.” | Coverage target not reached |
| Small population | All | “Population of [n] below 30; tested in full.” | Not applicable |
| Automated control | 1 plus general controls | “One instance per configuration variant, with reliance on the general controls tested at [ref].” | General controls found deficient: test more instances or treat as manual |
Evaluating results: exceptions, expansion, and the conclusion
Section 7 is written before testing and applied after it, and the rules below are the ones that keep the evaluation honest. Every exception is investigated for cause before it is counted, because a documentation lapse and a missing control get the same tally and different conclusions. The tolerable rate is fixed in advance and is not revised when the results arrive. Expansion is a decision with a trigger written in the memo, not a way of diluting a rate: the usual rule is that a single exception in a sample sized for zero triggers investigation and, if the cause is not isolated, a conclusion that the control is not effective, with expansion reserved for the case where the cause appears isolated and the auditor needs to confirm it. And a targeted selection is never projected; its exceptions are reported as facts about the items.
| Result | Action | Conclusion wording |
|---|---|---|
| No exceptions in a sample sized for zero | Conclude | “No exceptions in [n] of [N]; the control is concluded to be operating effectively for the period at [confidence] with a tolerable rate of [rate].” |
| One exception, cause isolated (a documented, one-off event) | Investigate and document the cause; expand by the original size if the memo provides for it; conclude on the combined result | “One exception in [n], attributed to [cause]; the sample was expanded by [n] with no further exceptions; the control is concluded effective with the exception reported as an observation.” |
| One exception, cause systemic (a design gap, a missing performer) | Stop; the control is not effective; assess the exposure with a targeted selection or full-population analytic | “The exception at item [x] arises from [systemic cause]; the control is concluded not to be operating effectively; the extent of exposure was assessed at [ref].” |
| Multiple exceptions | Do not expand; conclude ineffective; investigate causes; quantify exposure | “[n] exceptions in [N] exceed the tolerable rate; the control is concluded not effective for the period.” |
| Documentation shortfall without control failure | Record as an observation, not an exception, per section 6 | “Two items lacked evidence of the review date; the review was corroborated by [evidence]; recorded as a documentation observation.” |
| Item cannot be tested | Apply the replacement rule from section 5; a missing item that should exist is an exception | “Item [x] could not be located; treated as an exception per section 5 of the memo.” |
| Exceptions in a targeted selection | Report as facts; no rate, no projection | “3 of the 20 selected high-value items were not approved; no rate is inferred for the population.” |
The conclusion sentence is the last thing written and the first thing a reviewer reads, and it should be recognizable from section 8 of the memo: the same conclusion type, the same boundary, the same population. When the sentence a tester wants to write does not fit section 8, one of the two is wrong, and it is usually the sentence. The deficiency evaluation guide picks up where the evaluation ends, with what a concluded-ineffective control means for the report and, in a SOX context, for the classification.
The monetary-unit variant of the memo
When the objective is an amount rather than a rate, whether deposits are complete, whether a balance is fairly stated, whether the value of unapproved spend exceeds a threshold, the memo changes in three sections and stays the same in the rest. Section 1 names the conclusion as a misstatement amount with a tolerable figure. Section 4 records the monetary-unit parameters instead of a convention: the population value, the tolerable misstatement, the expected misstatement, the confidence level, and the resulting sampling interval and sample size, all of which mySampler calculates and records. Section 7 records how detected misstatements will be projected, by tainting for items below the interval and in full for items above it, and how the projected misstatement plus an allowance for sampling risk will be compared to the tolerable figure. MidState’s deposit test at E-1 was a memo of this kind: 70 deposits selected from $9,421,880 of receipts at a $134,600 interval, with every deposit above the interval selected automatically, so that the large deposits that would matter most to a completeness conclusion could not be missed. The rest of the memo, population and its reconciliation, selection mechanics with the seed, exception definition, and boundary, reads exactly as the rate version does, which is the point of a fixed template: the reviewer knows where to look regardless of the method.
Filling the memo from mySampler
The free mySampler tool on this site draws random, systematic, stratified, and monetary-unit samples in the browser and produces a documentation record with each draw. The record was designed around the memo, and the table shows where each piece of it goes, so that the selection section is copied rather than retyped and the seed that makes the draw reproducible is never lost.
| mySampler output | Memo section | Note |
|---|---|---|
| Population count, value, and the file or range loaded | 2. Population | The tool records what was loaded; the reconciliation to source is still the auditor’s, documented separately |
| Method chosen and its parameters (strata definitions, interval, start) | 3. Method | Copy the parameter line verbatim |
| Sample size and, for MUS, the parameter inputs | 4. Size | The rationale for the parameters is still written by the auditor |
| Random seed and timestamp | 5. Selection mechanics | The seed is what lets a reviewer reproduce the draw; record it every time |
| The selected item list with identifiers and stratum or interval | 5. Selection mechanics (listing reference) | Export it as the selection listing workpaper |
| Documentation record text | Attached to the memo | The record is evidence of the mechanics; the memo is evidence of the decisions |
Where the selection is judgmental, mySampler is not the tool; the criteria are applied in the population spreadsheet with a filter and the filter settings recorded in section 5. Either way the auditor performs the draw, on the auditor’s copy of the population, and the auditee’s role begins only when the item list exists. Functions that standardize on one tool for every draw gain something beyond convenience: the documentation record has the same shape in every file, so a reviewer or an assessor who has read one has read them all, and the seed, the parameters, and the timestamp are never the thing that is missing.
Worked example: MidState’s settlement reconciliation memo, completed
The memo below is the one the engagement lead wrote for the settlement reconciliation test in MidState Beverage’s FY27-01 route cash audit, before fieldwork, and it is reproduced with the section 9 entry added during testing. It is the memo that let the team say, when 14 of 60 reconciliations failed, exactly what that meant and what it did not.
Sampling memo. Engagement: FY27-01 Route cash handling. Test reference: C-3, settlement reconciliation. Prepared by: engagement lead, 12 January. Reviewed by: audit manager, 13 January.
1. Objective. Whether the daily settlement reconciliation was prepared and independently reviewed for each route settlement during the twelve months to 31 December. Conclusion sought: operating effectiveness rate for the population.
2. Population. All route settlements posted in the ERP route-accounting module for the twelve depots in the period: 37,400 settlements. Source: settlement table export, all depots, run 8 January by the analytics auditor with IT observing (C-1). Completeness: settlement count reconciled to route-day count from the dispatch system (300 routes, operating days per depot) with a difference of 0.3 percent explained by cancelled routes (C-1a); settlement value reconciled to route cash receipts in the GL with no unexplained difference (C-1b). The two acquired depots are on spreadsheets, not the ERP; their settlements (2,140) were obtained separately and added to the population (C-1c). Stratification: by depot, twelve strata.
3. Method. Stratified random, five per depot, because the objective is a population rate and the reconciliation is performed at depot level by different people with different staffing, so every depot must be represented.
4. Size. 60, five per depot. Basis: the function’s convention for high-frequency primary controls is 40 to 60; 60 chosen because this is the primary detective control over $31 million of cash, the area is rated high risk, and the FY26 loss went undetected for five months.
5. Selection mechanics. mySampler, stratified random, seed recorded in the documentation record at C-3a, drawn 12 January by the engagement lead. Items listed at C-3b. Depot clerks retrieved the reconciliation and handheld export for the listed settlements only; depots were not told the items in advance of the site visit. Replacement: a settlement with no reconciliation on file is an exception, not a replacement.
6. Exception. The reconciliation was not prepared; or was not signed by a reviewer; or the reviewer was the preparer; or the reviewer was the driver or handled the cash for that settlement (per the depot roster); or the review is dated more than two business days after the settlement. Not an exception: a missing date where the review is otherwise evidenced and the timing is corroborated by the scan timestamp; recorded as a documentation observation.
7. Evaluation. Tolerable rate 5 percent. Each exception investigated for cause before counting. One isolated exception: expand by 30 (stratified) and evaluate combined. Any systemic cause, or more than one exception: conclude the control not effective and assess exposure with the full-population override analysis at D-3. Projection: rate stated for the population with the stratification noted; no monetary projection from this test.
8. Boundary. Results support a conclusion on the operating effectiveness of the reconciliation across the twelve depots for the period. They support no conclusion on deposit completeness (E-1, monetary-unit sample) or on customer balances (F-3, targeted confirmations).
9. Changes. 2 February: at three depots the roster needed to test the independence attribute was not retained for the period; independence was tested from the handheld user log and the depot manager’s confirmation, approved by the audit manager (C-3c). No change to population, method, size, or exception definition. Result: 14 of 60 exceptions, at nine depots, all with a common cause (the reviewer was the preparer or had handled the cash, because depot staffing assumed a separation the depots did not have); the control was concluded not to be operating effectively; exposure assessed at D-3; expansion not performed because the cause was systemic.
Notice what the memo made possible. Section 2 caught the two acquired depots that the ERP export would have silently omitted. Section 6, written in January, meant that when a depot manager argued in March that a reviewer who “only counted the cash” was still independent, the answer was already on the page. Section 7 meant the team did not expand the sample in the hope of diluting fourteen failures into an acceptable rate; the cause was systemic, the memo said stop, and the exposure analysis took over. And section 9 recorded the one deviation with a date and an approval, which is exactly what a quality assessor looks for and almost never finds. The workpaper example shows the test workpaper this memo headed.
Where the memo sits in the file, and how a reviewer uses it
The memo is the first page of the test workpaper, ahead of the population reconciliation, the selection listing, the test matrix, and the conclusion, and it is the page the reviewer reads before anything else, because the review of a sample-based test is a review of the decisions before it is a review of the ticks. A reviewer who reads the memo first can check the test matrix against section 6 and the conclusion against section 8 in minutes; a reviewer who starts at the conclusion has to reconstruct the design from the results, which is slower and less reliable. The review itself follows a short sequence, and the questions below are the ones that find most problems.
| Reviewer question | Where the answer is | Red flag |
|---|---|---|
| Was the memo written and reviewed before testing? | Dates in the header | Memo dated after the first test date, or undated |
| Does the method fit the conclusion type? | Sections 1 and 3 | A rate conclusion from targeted items; a design test with a forty-item sample |
| Is the population complete? | Section 2 and the reconciliation reference | No reconciliation, or a reconciliation to the same report |
| Can the selection be reproduced? | Section 5: seed, tool, listing | No seed; items the auditee chose |
| Were exceptions applied as defined? | Section 6 against the test matrix | Items marked “minor” or “explained” outside the definition |
| Was the evaluation rule followed? | Section 7 against the conclusion | Expansion without a trigger; a rate from targeted items |
| Does the conclusion stay inside the boundary? | Section 8 against the conclusion sentence | A population claim the method does not support |
| Were deviations recorded and approved? | Section 9 | Population or size differs from the memo with no entry |
Retained review notes on the memo, cleared and kept in the file, are the evidence of supervision that the QAIP playbook asks for, and a function that reviews sampling memos this way finds that its file-level quality observations on sampling disappear within a cycle.
Common mistakes
| Mistake | What it looks like | Fix |
|---|---|---|
| Memo written after testing | Undated, or dated after the results | Date and reviewer initials before fieldwork; section 9 for anything that changed |
| Population equals the report | No reconciliation to anything independent | Section 2 names the reconciliation and its result |
| Size without a basis | “25 items were selected” | The convention, the parameters, or the criteria, plus the risk factors |
| No seed | A random draw nobody can reproduce | Record the seed and timestamp; mySampler’s record does it |
| Exception defined on the fly | The first questionable item starts a negotiation | Section 6 written in advance, with what does not count |
| Expansion as dilution | Sample doubled until the rate looks fine | Expansion only for an isolated cause, per a trigger in section 7 |
| Silent replacement | Unavailable items swapped for convenient ones | Replacement rule in section 5; missing items are usually exceptions |
| Projection from targeted items | A rate or an amount from the ten largest | Section 8 says no projection; the conclusion respects it |
| One memo for two objectives | Design and operation, or rate and amount, in one test | One memo per conclusion type |
| Changes unrecorded | The population shrank mid-test and nobody wrote it down | Section 9, dated, approved |
One page, nine sections, written first. The sampling memo is the cheapest document in the file to produce and the most expensive to be without, because every question a reviewer, an assessor, an external auditor, or an examiner will ask about a test is a question the memo answers in advance, in the auditor’s own words, on a date that proves the answer was a decision rather than a defense.
Related guides
- Audit sample sizes: 25, 40, 60
- Judgmental sampling that survives scrutiny
- Random sampling in Excel and Google Sheets
- Risk-directed sampling techniques
- mySampler (free tool)
- Testing information produced by the entity
- Audit evidence
- Audit workpaper example
- Test of design vs. operating effectiveness
- Control deficiency evaluation
- Audit work program
- The QAIP documentation kit
- Templates and downloads
- Start here
Leave a Reply