, ,

The Sampling Memo: A Template for Documenting Every Sampling Decision

Sampling decisions are made in the first hour of a test and defended in the last week of an external quality assessment, usually by someone who was not there. How was the population defined, and how do we know it was complete? Why forty and not twenty-five? How were the items chosen, and could anyone reproduce it? What counted as an exception, and what happened when one appeared? A workpaper that answers those questions in a paragraph at the top is a workpaper that survives review; one that leaves the reviewer to reverse-engineer the decisions from the test results is the reason so many file reviews end with an observation about sampling.

The sampling memo is one page, written before any item is tested, that records every decision in a fixed order: objective, population and its completeness, method, size and its rationale, selection mechanics, what an exception is, how results will be evaluated, and what the conclusion may say. This guide gives the template in full, annotates each section with what to write and what goes wrong, sets out the size rationales for the common cases, explains how exceptions are evaluated, shows how the outputs of the free mySampler tool drop into the memo, and works a completed memo from MidState Beverage’s route cash audit. It is the documentation companion to the sample size guide and the judgmental sampling guide.

In this guide

Why a memo, and why before testing

The Global Internal Audit Standards require engagement conclusions to rest on sufficient, reliable, relevant, and useful information and require the work to be documented so that a reviewer can understand what was done and why. Sampling is where those two requirements meet: the sufficiency of the evidence depends on how the sample was designed, and the design is invisible unless it is written down. The external audit standards, AU-C 530 and the PCAOB’s AS 2315, spell out the same elements, population, method, size, selection, evaluation, and reviewers trained on either set look for them in that order.

Writing the memo before testing matters for a reason that has nothing to do with paperwork. A size chosen before the results are known is a judgment; a size chosen afterward is a rationalization, and the two are indistinguishable on the page unless the memo carries a date and a reviewer’s initials from before fieldwork. The same is true of the definition of an exception: decided in advance, it is a standard; decided when the first questionable item appears, it is negotiable, and the auditee will negotiate it. The memo turns sampling from something the auditor did into something the auditor decided, which is what a reviewer is assessing.

ReaderWhat they look for in the memoWhat its absence costs
Engagement reviewerThat the design fits the objective and the size is justified before resultsReview notes, rework, and a test that may have to be redone
Quality assessorThe Standards’ documentation and evidence requirements met on the pageA file-level observation that repeats across engagements
External auditor considering relianceElements comparable to AS 2315 or AU-C 530, so the work can be usedRe-performance of the test at the company’s expense
Regulator or examinerPopulation completeness and independence of selectionA finding about the audit function
The auditor, a year laterWhat was done, to repeat or to defend itReconstructing decisions from memory, which fails

The sampling memo template

Sampling memo. Engagement: ________ Test reference: ________ Prepared by: ________ Date: ________ Reviewed by: ________ Date: ________

1. Test objective and the control or assertion tested. [One sentence: what the test will conclude on. “Whether the daily settlement reconciliation was prepared and independently reviewed for each route settlement in the period.”] Type of conclusion sought: [operating effectiveness rate / misstatement amount / targeted risk / design].

2. Population. Definition: [the set of items the conclusion will cover, with the period, the units, and any exclusions]. Source: [system, report name, parameters, extraction date, who extracted it]. Count and value: [n items; $ total]. Completeness check: [what the population was reconciled to, the result, and where the reconciliation is documented]. Stratification, if any: [strata and their counts and values].

3. Sampling method. [Random / systematic with a random start / stratified random / monetary-unit / targeted judgmental / full population.] Reason for the method: [why this method suits the objective and the population].

4. Sample size and rationale. Size: [n, or per stratum]. Basis: [the convention, the statistical parameters (confidence, tolerable rate, expected rate), or the judgmental criteria and coverage]. Risk assessment behind the size: [the control’s importance, the risk rating of the area, prior results, reliance placed on other testing].

5. Selection mechanics. Tool and settings: [mySampler / Excel formula / other; the random seed or start point recorded]. Selection performed by: [auditor name, date]. Items selected: [reference to the listing]. Independence: [the auditee’s role, if any, was limited to retrieving documents for items the auditor named]. Replacement rule: [what happens if a selected item cannot be tested: replaced by the next random item, or treated as an exception].

6. Definition of an exception. [What counts as a failure of the attribute tested, decided now: “the reconciliation was not prepared; or was prepared but not reviewed by someone other than the preparer; or the reviewer handled the cash for that settlement; or the review is undated or more than two business days after the settlement.”] What does not count: [documentation shortfalls that do not indicate control failure, and how they will be recorded].

7. Evaluation approach. Tolerable exception rate or amount: [n]. Action on exceptions: [each exception investigated for cause before counting; the sample expanded by [n] if [condition]; the control concluded ineffective if [condition]]. Projection: [whether and how results will be projected; none for targeted selections].

8. Conclusion boundary. [What the results will support and what they will not: population conclusion, targeted-risk conclusion, or design only.]

9. Changes after this memo. [Any deviation from sections 2 to 8 during testing, with the reason, the date, and reviewer approval. Blank means none.]

Section by section: what to write and what goes wrong

SectionWhat to writeWhat goes wrong
1. ObjectiveOne sentence naming the attribute or amount and the conclusion type; the conclusion type decides everything below it“Test invoices” as an objective; a design objective tested with a sample built for a rate
2. PopulationThe set the conclusion will cover, its source with parameters, count and value, and the reconciliation that shows it is completeA report accepted as the population without reconciling it to anything; two acquired depots missing because they are not on the ERP
3. MethodThe method and the reason it fits: random for a rate, monetary-unit for value, targeted for concentrated risk, full for small populations“Judgmental” with no criteria; random from a population where twenty items are 70 percent of the value
4. SizeThe number and its basis: the convention for the control’s frequency, the statistical parameters, or the judgmental coverage; the risk factors that pushed the size up or downA size chosen after the first ten items looked clean; no link between the area’s risk rating and the size
5. Selection mechanicsThe tool, the seed or start, who selected, the listing reference, the auditee’s role, the replacement ruleItems the auditee supplied; no seed, so the selection cannot be reproduced; unavailable items silently replaced by convenient ones
6. Exception definitionWhat fails the attribute, in the words a reviewer can apply; what does not count and how it is recordedDecided when the first questionable item appears; “minor” exceptions waved through without a rule
7. EvaluationTolerable rate, the investigate-before-counting rule, the expansion trigger, the projection method or its absenceExpanding the sample until the rate looks acceptable; projecting from a targeted selection
8. BoundaryWhat the conclusion may say, written nowA population conclusion drawn from the ten largest items
9. ChangesEvery deviation, dated and approvedSilent changes to the population or the size mid-test

The population section is where most memos fail and where the most effort belongs. A population is not a report; it is a report that has been reconciled to something independent of it, a control total, a ledger balance, a headcount, a prior-period count plus movements, and the reconciliation lives in the file. The IPE guide covers the completeness and accuracy testing that turns a report into a population, and the random sampling guide covers the tie-out mechanics in Excel and Sheets.

Size rationales for the common cases

Section 4 is a sentence or two, but the sentence has to come from somewhere, and the table gives the rationale for each common case in the form the memo needs. The conventions are those the sample size guide explains; they derive from attribute sampling at high confidence with a tolerable rate around ten percent and zero expected exceptions, which is why one exception changes everything.

CaseSizeRationale to write in section 4Adjust upward when
Annual control1“The control operates once; the single instance is tested in full with re-performance.”Never; re-perform instead
Quarterly2, or all 4 for key controls“Two of four instances per the function’s convention for quarterly controls; all four because the control is key to [assertion].”Key control; prior exceptions
Monthly2 to 4“Three of twelve per the convention for monthly controls, increased from two because the area is rated high risk.”High risk; new control; prior exceptions
Weekly5 to 10“Eight of 52 per the convention, the upper end because the control was new in the period.”New control; turnover in the operator
Daily20 to 25, up to 40“Twenty-five of approximately 250 per the convention for daily controls; 40 where the control is the primary control over [risk].”Primary control; high risk; no compensating control
Many times daily25 to 40, up to 60“Forty of 37,400 per the convention for high-frequency controls, stratified by depot so that every location is represented; 60 where the control is primary and the area high risk.”Primary control; high risk; heterogeneous population
Monetary-unit (value conclusion)Calculated“Sample size of [n] from the MUS parameters: population value [$], tolerable misstatement [$], expected misstatement [$], confidence [percent]; interval [$].”Lower tolerable misstatement; higher expected error
Targeted judgmentalBy criteria“All items meeting [criteria], [n] items covering [percent] of value; no population rate will be concluded.”Coverage target not reached
Small populationAll“Population of [n] below 30; tested in full.”Not applicable
Automated control1 plus general controls“One instance per configuration variant, with reliance on the general controls tested at [ref].”General controls found deficient: test more instances or treat as manual

Evaluating results: exceptions, expansion, and the conclusion

Section 7 is written before testing and applied after it, and the rules below are the ones that keep the evaluation honest. Every exception is investigated for cause before it is counted, because a documentation lapse and a missing control get the same tally and different conclusions. The tolerable rate is fixed in advance and is not revised when the results arrive. Expansion is a decision with a trigger written in the memo, not a way of diluting a rate: the usual rule is that a single exception in a sample sized for zero triggers investigation and, if the cause is not isolated, a conclusion that the control is not effective, with expansion reserved for the case where the cause appears isolated and the auditor needs to confirm it. And a targeted selection is never projected; its exceptions are reported as facts about the items.

ResultActionConclusion wording
No exceptions in a sample sized for zeroConclude“No exceptions in [n] of [N]; the control is concluded to be operating effectively for the period at [confidence] with a tolerable rate of [rate].”
One exception, cause isolated (a documented, one-off event)Investigate and document the cause; expand by the original size if the memo provides for it; conclude on the combined result“One exception in [n], attributed to [cause]; the sample was expanded by [n] with no further exceptions; the control is concluded effective with the exception reported as an observation.”
One exception, cause systemic (a design gap, a missing performer)Stop; the control is not effective; assess the exposure with a targeted selection or full-population analytic“The exception at item [x] arises from [systemic cause]; the control is concluded not to be operating effectively; the extent of exposure was assessed at [ref].”
Multiple exceptionsDo not expand; conclude ineffective; investigate causes; quantify exposure“[n] exceptions in [N] exceed the tolerable rate; the control is concluded not effective for the period.”
Documentation shortfall without control failureRecord as an observation, not an exception, per section 6“Two items lacked evidence of the review date; the review was corroborated by [evidence]; recorded as a documentation observation.”
Item cannot be testedApply the replacement rule from section 5; a missing item that should exist is an exception“Item [x] could not be located; treated as an exception per section 5 of the memo.”
Exceptions in a targeted selectionReport as facts; no rate, no projection“3 of the 20 selected high-value items were not approved; no rate is inferred for the population.”

The conclusion sentence is the last thing written and the first thing a reviewer reads, and it should be recognizable from section 8 of the memo: the same conclusion type, the same boundary, the same population. When the sentence a tester wants to write does not fit section 8, one of the two is wrong, and it is usually the sentence. The deficiency evaluation guide picks up where the evaluation ends, with what a concluded-ineffective control means for the report and, in a SOX context, for the classification.

The monetary-unit variant of the memo

When the objective is an amount rather than a rate, whether deposits are complete, whether a balance is fairly stated, whether the value of unapproved spend exceeds a threshold, the memo changes in three sections and stays the same in the rest. Section 1 names the conclusion as a misstatement amount with a tolerable figure. Section 4 records the monetary-unit parameters instead of a convention: the population value, the tolerable misstatement, the expected misstatement, the confidence level, and the resulting sampling interval and sample size, all of which mySampler calculates and records. Section 7 records how detected misstatements will be projected, by tainting for items below the interval and in full for items above it, and how the projected misstatement plus an allowance for sampling risk will be compared to the tolerable figure. MidState’s deposit test at E-1 was a memo of this kind: 70 deposits selected from $9,421,880 of receipts at a $134,600 interval, with every deposit above the interval selected automatically, so that the large deposits that would matter most to a completeness conclusion could not be missed. The rest of the memo, population and its reconciliation, selection mechanics with the seed, exception definition, and boundary, reads exactly as the rate version does, which is the point of a fixed template: the reviewer knows where to look regardless of the method.

Filling the memo from mySampler

The free mySampler tool on this site draws random, systematic, stratified, and monetary-unit samples in the browser and produces a documentation record with each draw. The record was designed around the memo, and the table shows where each piece of it goes, so that the selection section is copied rather than retyped and the seed that makes the draw reproducible is never lost.

mySampler outputMemo sectionNote
Population count, value, and the file or range loaded2. PopulationThe tool records what was loaded; the reconciliation to source is still the auditor’s, documented separately
Method chosen and its parameters (strata definitions, interval, start)3. MethodCopy the parameter line verbatim
Sample size and, for MUS, the parameter inputs4. SizeThe rationale for the parameters is still written by the auditor
Random seed and timestamp5. Selection mechanicsThe seed is what lets a reviewer reproduce the draw; record it every time
The selected item list with identifiers and stratum or interval5. Selection mechanics (listing reference)Export it as the selection listing workpaper
Documentation record textAttached to the memoThe record is evidence of the mechanics; the memo is evidence of the decisions

Where the selection is judgmental, mySampler is not the tool; the criteria are applied in the population spreadsheet with a filter and the filter settings recorded in section 5. Either way the auditor performs the draw, on the auditor’s copy of the population, and the auditee’s role begins only when the item list exists. Functions that standardize on one tool for every draw gain something beyond convenience: the documentation record has the same shape in every file, so a reviewer or an assessor who has read one has read them all, and the seed, the parameters, and the timestamp are never the thing that is missing.

Worked example: MidState’s settlement reconciliation memo, completed

The memo below is the one the engagement lead wrote for the settlement reconciliation test in MidState Beverage’s FY27-01 route cash audit, before fieldwork, and it is reproduced with the section 9 entry added during testing. It is the memo that let the team say, when 14 of 60 reconciliations failed, exactly what that meant and what it did not.

Sampling memo. Engagement: FY27-01 Route cash handling. Test reference: C-3, settlement reconciliation. Prepared by: engagement lead, 12 January. Reviewed by: audit manager, 13 January.

1. Objective. Whether the daily settlement reconciliation was prepared and independently reviewed for each route settlement during the twelve months to 31 December. Conclusion sought: operating effectiveness rate for the population.

2. Population. All route settlements posted in the ERP route-accounting module for the twelve depots in the period: 37,400 settlements. Source: settlement table export, all depots, run 8 January by the analytics auditor with IT observing (C-1). Completeness: settlement count reconciled to route-day count from the dispatch system (300 routes, operating days per depot) with a difference of 0.3 percent explained by cancelled routes (C-1a); settlement value reconciled to route cash receipts in the GL with no unexplained difference (C-1b). The two acquired depots are on spreadsheets, not the ERP; their settlements (2,140) were obtained separately and added to the population (C-1c). Stratification: by depot, twelve strata.

3. Method. Stratified random, five per depot, because the objective is a population rate and the reconciliation is performed at depot level by different people with different staffing, so every depot must be represented.

4. Size. 60, five per depot. Basis: the function’s convention for high-frequency primary controls is 40 to 60; 60 chosen because this is the primary detective control over $31 million of cash, the area is rated high risk, and the FY26 loss went undetected for five months.

5. Selection mechanics. mySampler, stratified random, seed recorded in the documentation record at C-3a, drawn 12 January by the engagement lead. Items listed at C-3b. Depot clerks retrieved the reconciliation and handheld export for the listed settlements only; depots were not told the items in advance of the site visit. Replacement: a settlement with no reconciliation on file is an exception, not a replacement.

6. Exception. The reconciliation was not prepared; or was not signed by a reviewer; or the reviewer was the preparer; or the reviewer was the driver or handled the cash for that settlement (per the depot roster); or the review is dated more than two business days after the settlement. Not an exception: a missing date where the review is otherwise evidenced and the timing is corroborated by the scan timestamp; recorded as a documentation observation.

7. Evaluation. Tolerable rate 5 percent. Each exception investigated for cause before counting. One isolated exception: expand by 30 (stratified) and evaluate combined. Any systemic cause, or more than one exception: conclude the control not effective and assess exposure with the full-population override analysis at D-3. Projection: rate stated for the population with the stratification noted; no monetary projection from this test.

8. Boundary. Results support a conclusion on the operating effectiveness of the reconciliation across the twelve depots for the period. They support no conclusion on deposit completeness (E-1, monetary-unit sample) or on customer balances (F-3, targeted confirmations).

9. Changes. 2 February: at three depots the roster needed to test the independence attribute was not retained for the period; independence was tested from the handheld user log and the depot manager’s confirmation, approved by the audit manager (C-3c). No change to population, method, size, or exception definition. Result: 14 of 60 exceptions, at nine depots, all with a common cause (the reviewer was the preparer or had handled the cash, because depot staffing assumed a separation the depots did not have); the control was concluded not to be operating effectively; exposure assessed at D-3; expansion not performed because the cause was systemic.

Notice what the memo made possible. Section 2 caught the two acquired depots that the ERP export would have silently omitted. Section 6, written in January, meant that when a depot manager argued in March that a reviewer who “only counted the cash” was still independent, the answer was already on the page. Section 7 meant the team did not expand the sample in the hope of diluting fourteen failures into an acceptable rate; the cause was systemic, the memo said stop, and the exposure analysis took over. And section 9 recorded the one deviation with a date and an approval, which is exactly what a quality assessor looks for and almost never finds. The workpaper example shows the test workpaper this memo headed.

Where the memo sits in the file, and how a reviewer uses it

The memo is the first page of the test workpaper, ahead of the population reconciliation, the selection listing, the test matrix, and the conclusion, and it is the page the reviewer reads before anything else, because the review of a sample-based test is a review of the decisions before it is a review of the ticks. A reviewer who reads the memo first can check the test matrix against section 6 and the conclusion against section 8 in minutes; a reviewer who starts at the conclusion has to reconstruct the design from the results, which is slower and less reliable. The review itself follows a short sequence, and the questions below are the ones that find most problems.

Reviewer questionWhere the answer isRed flag
Was the memo written and reviewed before testing?Dates in the headerMemo dated after the first test date, or undated
Does the method fit the conclusion type?Sections 1 and 3A rate conclusion from targeted items; a design test with a forty-item sample
Is the population complete?Section 2 and the reconciliation referenceNo reconciliation, or a reconciliation to the same report
Can the selection be reproduced?Section 5: seed, tool, listingNo seed; items the auditee chose
Were exceptions applied as defined?Section 6 against the test matrixItems marked “minor” or “explained” outside the definition
Was the evaluation rule followed?Section 7 against the conclusionExpansion without a trigger; a rate from targeted items
Does the conclusion stay inside the boundary?Section 8 against the conclusion sentenceA population claim the method does not support
Were deviations recorded and approved?Section 9Population or size differs from the memo with no entry

Retained review notes on the memo, cleared and kept in the file, are the evidence of supervision that the QAIP playbook asks for, and a function that reviews sampling memos this way finds that its file-level quality observations on sampling disappear within a cycle.

Common mistakes

MistakeWhat it looks likeFix
Memo written after testingUndated, or dated after the resultsDate and reviewer initials before fieldwork; section 9 for anything that changed
Population equals the reportNo reconciliation to anything independentSection 2 names the reconciliation and its result
Size without a basis“25 items were selected”The convention, the parameters, or the criteria, plus the risk factors
No seedA random draw nobody can reproduceRecord the seed and timestamp; mySampler’s record does it
Exception defined on the flyThe first questionable item starts a negotiationSection 6 written in advance, with what does not count
Expansion as dilutionSample doubled until the rate looks fineExpansion only for an isolated cause, per a trigger in section 7
Silent replacementUnavailable items swapped for convenient onesReplacement rule in section 5; missing items are usually exceptions
Projection from targeted itemsA rate or an amount from the ten largestSection 8 says no projection; the conclusion respects it
One memo for two objectivesDesign and operation, or rate and amount, in one testOne memo per conclusion type
Changes unrecordedThe population shrank mid-test and nobody wrote it downSection 9, dated, approved

One page, nine sections, written first. The sampling memo is the cheapest document in the file to produce and the most expensive to be without, because every question a reviewer, an assessor, an external auditor, or an examiner will ask about a test is a question the memo answers in advance, in the auditor’s own words, on a date that proves the answer was a decision rather than a defense.

Related guides

Comments

Leave a Reply

Discover more from internalauditguide.com

Subscribe now to keep reading and get access to the full archive.

Continue reading