,

Judgmental Sampling That Survives Scrutiny: Documenting Non-Statistical Selections

Every reviewer has read the sentence “a sample of 25 invoices was selected for testing” and wondered how. Randomly, from a defined population, in which case the results say something about the population? By picking the largest, in which case they say something about the largest? Or by taking whatever the auditee handed over, in which case they say nothing at all? The workpaper rarely answers, and the difference between the three is the difference between a conclusion that survives an external quality assessment, a regulator’s file review, or a courtroom, and one that does not. Judgmental selection is legitimate and often the best choice; undocumented judgmental selection is indistinguishable from convenience, and that is what reviewers assume it was.

This guide is about the non-statistical selection that internal auditors use every day and rarely document properly. It separates risk-directed judgmental selection from haphazard and convenience selection, sets out when judgment beats a random sample, gives the seven elements a judgmental selection needs in the workpaper, contrasts the language that distinguishes deliberate selection from grabbing, states plainly what a judgmental selection can and cannot support in a conclusion, and provides six rationale paragraphs written to be copied into a workpaper and edited. MidState Beverage’s route cash audit, which combined a statistical sample with targeted selections, is the worked example. The random sampling guide covers the statistical side; the sampling memo template is where the documentation described here lives.

In this guide

Three kinds of selection, and what each one supports

The external audit standards draw a line that internal auditors should borrow. Audit sampling, in AICPA AU-C 530 and the PCAOB’s AS 2315, means applying a procedure to less than all of a population in a way that gives every item a chance of selection, so that the results can be projected to the population; the PCAOB’s amended AS 2315 takes effect on 15 December 2026 and keeps the distinction. Selecting specific items, the high-value ones, the unusual ones, the ones a risk assessment points at, is a different procedure that supports conclusions about the items selected and about the risks they represent, but not a projection to the whole. Both are legitimate. The third kind, haphazard or convenience selection, is what happens when the auditor takes items without a defined method, and it supports nothing, because nobody can say what the items represent.

KindHow items are chosenWhat the results supportWhen it is the right choiceWhat a reviewer asks
Statistical (random, systematic, monetary-unit)From a defined, reconciled population, by a method that gives each item, or each dollar, a known chanceA conclusion about the population: an exception rate, a projected misstatement, with a stated confidenceThe objective is a population-level conclusion on operating effectiveness or on the amount of errorPopulation definition and completeness; method; size rationale; evaluation
Judgmental, risk-directedSpecific items chosen against explicit criteria: value, risk attributes, timing, source, prior historyA conclusion about the items tested and about the risk the criteria targeted: whether the controls held where the exposure was highest, whether a suspected pattern existsConcentrated value or risk; small populations; testing design; investigating a known problem; regulatory expectations for specific itemsThe criteria, the coverage achieved, and whether the conclusion stayed within what the items support
Haphazard or convenienceNo defined method: what was on top, what the auditee provided, what was easy to obtainNothing beyond the items themselves; no basis to say what they representNever, for a conclusion; at most for gaining an initial understanding before a real selection“How were these chosen?” — and the answer ends the discussion

The word “haphazard” has a technical meaning in the sampling literature, selection without conscious bias that is treated as approximately random, and internal auditors should avoid the term because reviewers read it as “no method”. If a selection was meant to approximate random, use a random method; the mySampler tool draws one in seconds. If it was meant to target risk, say so and say how. The middle ground, where the auditor meant to be random but picked from the top of the file, is the one that fails review.

When judgmental selection is the right call

Judgmental selection is not the fallback for auditors who cannot be bothered with a random sample. In the situations below it is the better method, and a random sample would be the weaker choice, because the objective is not a population-level rate but a question about where the risk sits.

SituationWhy judgment beats randomSelection criteria that workConclusion the selection supports
Value is concentratedTwenty items are 70 percent of the value; a random sample of 25 from 4,000 would probably miss them allAll items above a threshold, or the top n by value, plus a random sample of the remainder if a population conclusion is also neededThe controls operated over most of the value at risk
Small populationFewer than 30 items: test all, or all the ones that matter100 percent, or all items meeting a risk attributeA conclusion on the whole population, because the whole population was tested
Testing design or understandingOne or two transactions walked end to end show whether the control would work; randomness adds nothingA typical item and an atypical one (an exception, a manual override, a period-end entry)Design adequacy; not operating effectiveness
Investigating a known problemA prior finding, a complaint, an analytic anomaly: the question is whether the pattern exists, not how oftenItems with the anomaly’s attributes: the depot, the user, the date range, the vendorWhether the pattern is confirmed and its extent within the criteria
Regulatory or contractual specific itemsExaminers and contracts often specify what must be looked atThe items named: high-risk customers, related-party transactions, the largest exposuresCompliance for the specified items
Population cannot be defined reliablyWithout a complete population a random sample has no meaning; targeted selection from what exists is honest about the limitationItems from the available records with the limitation statedA limited conclusion with the limitation disclosed; and a finding about the missing population
Timing riskPeriod-end, weekends, holidays, and cutover dates carry more risk than the rest of the yearAll items in the risk windowWhether the control held at the point of highest risk
Unusual attributesRound amounts, duplicates, one-time vendors, entries by unusual users, split transactionsAnalytics-driven lists of items with the attribute; see the journal entry analytics guideWhether the attribute indicates a control gap or fraud; extent within the list

The pattern across the table is that judgmental selection answers “where is the risk and did the control hold there”, while statistical sampling answers “how often does the control fail”. Most engagements need both answers, which is why the strongest designs, discussed under conclusions below, combine a targeted selection of the items that matter most with a random sample of the rest. The risk-directed sampling guide covers the design choices in depth.

The seven elements of a judgmental selection that survives review

A judgmental selection is defensible when a reviewer who was not there can reconstruct why each item was chosen and what the items together represent. That takes seven elements in the workpaper, written before the testing rather than after it, because criteria written after the results are known are criteria fitted to the answer.

ElementWhat it recordsWhy a reviewer needs it
1. ObjectiveWhat question the selection is meant to answer: design, operation at the highest-risk points, existence of a suspected patternThe objective determines whether judgment is the right method and what the conclusion may say
2. PopulationThe full set the items were chosen from, its source, the count and value, and the completeness check performedJudgmental selection still needs a population; “the ten largest” means nothing without the list they were largest in
3. Selection criteriaThe explicit rules: thresholds, attributes, date windows, risk scores, in the order appliedCriteria are what separate risk-directed selection from convenience; they must be reproducible
4. The items and their reason codesEach item listed with which criterion selected itA reviewer can see that the criteria were actually applied and nothing was added by hand without a reason
5. Coverage achievedThe items as a share of the population by count and by value; the share of the risk attribute covered“Twenty items covering 71 percent of value” is a statement of what the test reached; “twenty items” alone is not
6. Independence of selectionConfirmation that the auditor selected from the population, not from items the auditee proposed or pre-assembledAuditee-selected items are the most common way a judgmental selection becomes worthless
7. Conclusion boundaryA statement, written before testing, of what the results will and will not supportPrevents the finished workpaper from claiming a population conclusion the method cannot give

Element six deserves its own warning. Process owners are helpful people, and “here are some examples for you” is the sentence that has undone more judgmental selections than any other. The auditor obtains the population, applies the criteria, and tells the auditee which items are needed; the auditee retrieves them. Where the population itself comes from the auditee, as it usually does, the completeness check in element two is what makes the selection independent, and the IPE guide covers how to perform it. The workpaper example shows the seven elements laid out on a page.

Language: deliberate selection versus grabbing

The same twenty items can be described in a way that reads as method or in a way that reads as convenience, and reviewers respond to the words. The table pairs the phrasing that fails review with the phrasing that passes; the difference in every row is that the passing version names the population, the criteria, and the coverage.

Reads as grabbingReads as method
A sample of 20 invoices was selected for testing.From the 4,212 invoices posted in the period (AP register, reconciled to the GL, workpaper C-2), the 20 highest-value invoices from vendors created in the period were selected, covering $1.84 million or 71 percent of new-vendor spend.
Judgmentally selected 10 journal entries.All 10 manual journal entries above $100,000 posted in the last three business days of the quarter were selected (population: 1,140 manual entries in the quarter, JE analytics workpaper D-1); these represent 58 percent of period-end manual entry value.
Reviewed a selection of user accounts.All 14 accounts holding the ERP system administrator role and all 9 accounts with vendor-master and payment-release access were selected from the role export dated 30 June (population 612 active accounts); together they are every account able to change or release a payment without a second person.
Tested several reconciliations.The bank reconciliations for the two accounts with the highest transaction volume and the one account that showed unreconciled items in the prior audit were selected for all three months of the quarter (9 reconciliations of 27 performed).
Obtained examples of contracts from the procurement team.From the contract register (247 contracts active at 30 June, reconciled to the legal department’s list), the auditor selected the 12 contracts above $500,000 in annual value and the 5 contracts with related-party counterparties, covering 64 percent of contracted spend.
A number of depots were visited.Four of twelve depots were selected for site visits: the two with the highest override rates in the full-population analysis and the two acquired depots not yet on the ERP; together they hold 41 percent of route cash volume.

Three habits produce the right-hand column. Always state the population with its source and reconciliation reference before describing the selection. Always use a number for the criterion, never a word like “large” or “significant”. And always finish with the coverage, so that the reader knows what share of the thing at risk the test reached. The workpaper best practices guide treats this as a documentation standard rather than a style preference, which is what it is.

What you can and cannot conclude

A judgmental selection supports three kinds of conclusion and cannot support two others, and the workpaper’s conclusion sentence must stay on the right side of the line. It can conclude on the items tested: these twenty invoices were properly approved, or three were not. It can conclude on the risk the criteria targeted: the controls operated over the highest-value new-vendor spend, or they did not at the point of highest exposure. And it can conclude on design, where the objective was design. It cannot state an exception rate for the population, because the items were not chosen in a way that represents it, and it cannot project a misstatement, for the same reason. “3 of 20 exceptions, 15 percent” is a fact about the twenty; “15 percent of invoices are unapproved” is a claim the method does not support, and reviewers know the difference.

Conclusion typeSupported by judgmental selection?Example wording
On the items testedYes“All 20 selected invoices were approved by an authorized approver before payment.”
On the targeted riskYes, within the criteria“Approval controls operated over the 71 percent of new-vendor spend represented by the selection; the control was not tested for lower-value invoices.”
On designYes, where that was the objective“The walkthrough of two transactions confirmed the control is designed to require a second approver above $10,000.”
Population exception rateNoNot: “15 percent of invoices lacked approval.” Instead: “3 of the 20 highest-value new-vendor invoices lacked approval; the extent across the remaining population was not tested and is addressed by the random sample at C-4.”
Projected misstatementNoNot: “unapproved spend is estimated at $410,000.” Instead: “unapproved spend within the selection totaled $61,200; no projection is made.”
Operating effectiveness of the control across the periodOnly when combined with a statistical sample of the remainder, or when the selection is the full population“The 20 targeted items and the random sample of 40 from the remaining population together support the conclusion that the control operated effectively; see C-4 for the sample evaluation.”

The combined design in the last row is the professional answer to most engagements. Take the items that matter most by value or risk in full, then draw a random sample from what remains, evaluate the two parts separately, and conclude on the population from the sample and on the risk from the targeted items. Monetary-unit sampling does something similar automatically, since large items are almost certain to be selected, which is why the sample size guide recommends it for value-weighted populations. What the combined design never does is mix the exceptions from the two parts into one rate; the three targeted exceptions and the one random exception are reported as what they are.

Six rationale paragraphs to copy

Each paragraph below is a complete selection rationale for a workpaper, with the seven elements in order. Replace the bracketed items, keep the structure, and write it before the testing starts.

1. High-value coverage. Objective: to test whether the approval control operated over the invoices carrying most of the value at risk. Population: [n] invoices totaling [$] posted between [dates], per the AP register at [ref], reconciled to the GL at [ref] with no unreconciled differences. Criteria: all invoices above [$threshold], applied in descending value order. Items: [n] invoices listed at [ref] with the criterion recorded against each. Coverage: [n] of [n] invoices, [percent] of value. Independence: the auditor extracted the register and applied the criteria; the auditee supplied documents for the listed items only. Conclusion boundary: results support a conclusion on the selected items and on the control’s operation over [percent] of value; no exception rate or projection to the population is made; population-level conclusions rest on the random sample at [ref].

2. New or changed master data. Objective: to test whether vendor onboarding controls operated for vendors created or changed in the period, where the fraud and error risk concentrates. Population: [n] vendor master records created or with bank details changed between [dates], per the change log at [ref], reconciled to the vendor master count. Criteria: all records with a bank account change, plus all new vendors with payments above [$] in the period. Items: [n] at [ref]. Coverage: 100 percent of bank changes; [percent] of new-vendor spend. Independence: extracted by the auditor from the system. Conclusion boundary: supports a conclusion on onboarding and change controls for the items and the targeted risk; not an exception rate for all vendor changes.

3. Period-end manual journals. Objective: to test whether manual journal entries at period end, the point of highest misstatement risk, were reviewed and supported. Population: [n] manual entries in the quarter per the JE extract at [ref], reconciled to the GL posting total. Criteria: entries above [$] posted in the last [n] business days, entries by users outside the accounting team, and entries with round amounts above [$]. Items: [n] at [ref], each with the criterion. Coverage: [percent] of period-end manual entry value. Independence: extracted and filtered by the auditor. Conclusion boundary: supports conclusions on the items and on the period-end risk; the operating effectiveness of JE review across the quarter is addressed by the sample at [ref].

4. Privileged access. Objective: to test whether the accounts able to bypass segregation of duties are restricted and reviewed. Population: [n] active accounts in [system] per the role export dated [date] at [ref], reconciled to the HR active headcount with [n] differences investigated at [ref]. Criteria: all accounts holding [administrator role] and all accounts holding both [conflicting role A] and [conflicting role B]. Items: [n] accounts at [ref]. Coverage: 100 percent of the privileged population. Independence: extracted by the auditor with IT observing. Conclusion boundary: supports a conclusion on the whole privileged population, since all of it was tested; supports no conclusion about the appropriateness of standard-user access, which is addressed at [ref].

5. Follow-on from an exception. Objective: to determine whether the exception found at [ref] is isolated or indicates a pattern. Population: [n] items sharing the exception’s attributes ([depot, user, vendor, date range]) per [ref]. Criteria: all items in the population, or where the population exceeds [n], the [n] most recent and the [n] highest-value. Items: [n] at [ref]. Coverage: [percent] of the attribute population. Independence: defined by the auditor from the original exception. Conclusion boundary: supports a conclusion on whether the pattern exists within the attribute population and its extent there; supports no conclusion about items outside the attributes.

6. Small population tested in full. Objective: to conclude on the operating effectiveness of [control] for the period. Population: [n] instances of the control in the period per [ref], reconciled to [source]. Criteria: none; the population is below [30] items and was tested in full. Items: all [n] at [ref]. Coverage: 100 percent. Independence: population obtained by the auditor from [system]. Conclusion boundary: supports a conclusion on the population, since the population was tested; any exception is an exception in the population, not a sample result.

Worked example: MidState’s route cash confirmations

MidState Beverage’s FY27-01 route cash audit used both kinds of selection deliberately, and the workpapers say which was which. For the settlement reconciliations, the objective was a population-level conclusion on whether the daily control operated, so the team drew a stratified random sample: 60 reconciliations, five per depot, from the year’s 37,400 settlements, evaluated as a sample; 14 failed, and the conclusion on operating effectiveness rested on that sample and the full-population override analysis. For customer balances, the objective was different: the FY26 diversion had run for five months because the affected customers received no statements and never noticed, so the question was whether a similar loss could be sitting undetected now, at the customers where it would most likely sit. That is a judgmental question, and the selection rationale in the workpaper read as follows.

Objective: to determine whether cash-paying customers who receive no monthly statement carry balances that differ from the ledger, indicating diverted or misapplied payments. Population: 1,130 customers flagged in the customer master as receiving no statement (workpaper F-2, reconciled to the 4,200 cash-paying customers at F-1), with ledger balances totaling $9,421,880 at 30 April. Criteria: the 12 customers with the largest ledger balances, plus the 8 customers with the highest cash volume at the four depots with the highest override rates in the full-population analysis (D-3). Items: 20 customers listed at F-3 with the criterion against each. Coverage: 20 of 1,130 unstated customers; 38 percent of the unstated ledger balance; four depots holding 41 percent of route cash volume. Independence: the auditor selected from the customer master extract; depot staff were not told which customers would be confirmed until the confirmations were sent by the audit team directly. Conclusion boundary: results support a conclusion on the 20 customers and on the targeted risk, whether large unstated balances at high-override depots are misstated; they support no exception rate for the 1,130, and the population-level question of statement coverage is a finding in its own right (F5) regardless of confirmation results. A monetary-unit sample of 70 deposits (workpaper E-1) separately supports the conclusion on deposit completeness across the population.

The rationale did its job in two directions. When the confirmations came back, the team could say exactly what the results meant: the 20 balances agreed within tolerance, which showed no diversion at the highest-risk points and did not show that the other 1,110 customers were fine; the finding that 1,130 customers received no statement stood on its own as a control gap. And when the COO argued at the closing meeting that the clean confirmations proved the statement gap was harmless, the conclusion boundary, written weeks earlier, was the answer: the confirmations were selected to find a loss where it was likeliest, not to measure the population, and a control that would let a loss run undetected at any of 1,130 customers is a control gap whether or not a loss was running that month. The root cause guide covers how the finding was written; the substantive testing guide covers the confirmation procedure itself.

How external auditors and examiners read a judgmental selection

Two outside readers judge internal audit’s selections, and both apply the distinction this guide rests on. An external auditor deciding whether to use internal audit’s work under PCAOB AS 2605 or the AICPA equivalent asks whether the testing supports the conclusion drawn from it; a judgmental selection presented as an operating-effectiveness conclusion will be discarded and re-performed, while the same selection documented with its criteria, coverage, and boundary, alongside a random sample of the remainder, is work the external auditor can use. Bank and insurance examiners read files the same way, and an examiner who finds “management provided examples” in a selection rationale reads it as a finding about the audit function rather than about the process. The regulator expectations guide covers what examiners look for in workpapers; the short version is that they look for exactly the seven elements above.

There is one more reader: the function’s own quality program. A self-assessment or external quality assessment reviews engagement files against the Standards’ evidence requirements, and undocumented selection is among the most common file-level observations assessors raise, because it is visible on every page and easy to fix. A function that adopts the seven elements as a documentation standard, checks them at the fieldwork QC checkpoint described in the QAIP kit, and puts the rationale in the sampling memo removes the observation before it is made.

Common mistakes

MistakeWhat it looks likeFix
“Judgmental” as a synonym for undocumentedThe word appears in the workpaper with no criteriaCriteria, coverage, and boundary written before testing
Auditee-provided examples“Management provided ten contracts for review”The auditor obtains the population and selects; the auditee retrieves
Criteria written after the resultsSelection rationale added at review time to fit what was testedRationale dated before fieldwork; the reviewer checks the date
Rates from targeted items“3 of 20, a 15 percent exception rate”Report the exceptions as facts about the items; rates come from samples
No population“The largest items were selected” with no list they were largest inPopulation, source, count, value, completeness check, every time
Coverage unstatedTwenty items, and no idea what share of valueCount and value coverage in the rationale and the conclusion
Judgment where a rate was neededA SOX operating-effectiveness conclusion drawn from the ten largest transactionsRandom sample of the population, or targeted items plus a sample of the remainder
Random where judgment was neededA random 25 from 4,000 invoices that misses every large one, and a clean conclusion over 4 percent of valueTargeted high-value items first; then the sample
Mixing the two sets of exceptionsTargeted and sampled exceptions combined into one rateEvaluate separately; report separately
The word “haphazard”Used to mean “we picked some”Use a random method or a documented judgmental one; never the word

Judgmental selection is the auditor’s most powerful tool for putting testing where the risk is, and the least respected, because it is so rarely written down properly. Seven elements, written before the work, and a conclusion that stays within what the items support: that is the whole discipline, and a selection documented that way survives any reviewer, because the reviewer can see exactly what was done and exactly what it means.

Related guides

Comments

Leave a Reply

Discover more from internalauditguide.com

Subscribe now to keep reading and get access to the full archive.

Continue reading