Every reviewer has read the sentence “a sample of 25 invoices was selected for testing” and wondered how. Randomly, from a defined population, in which case the results say something about the population? By picking the largest, in which case they say something about the largest? Or by taking whatever the auditee handed over, in which case they say nothing at all? The workpaper rarely answers, and the difference between the three is the difference between a conclusion that survives an external quality assessment, a regulator’s file review, or a courtroom, and one that does not. Judgmental selection is legitimate and often the best choice; undocumented judgmental selection is indistinguishable from convenience, and that is what reviewers assume it was.
This guide is about the non-statistical selection that internal auditors use every day and rarely document properly. It separates risk-directed judgmental selection from haphazard and convenience selection, sets out when judgment beats a random sample, gives the seven elements a judgmental selection needs in the workpaper, contrasts the language that distinguishes deliberate selection from grabbing, states plainly what a judgmental selection can and cannot support in a conclusion, and provides six rationale paragraphs written to be copied into a workpaper and edited. MidState Beverage’s route cash audit, which combined a statistical sample with targeted selections, is the worked example. The random sampling guide covers the statistical side; the sampling memo template is where the documentation described here lives.
In this guide
- Three kinds of selection, and what each one supports
- When judgmental selection is the right call
- The seven elements of a judgmental selection that survives review
- Language: deliberate selection versus grabbing
- What you can and cannot conclude
- Six rationale paragraphs to copy
- Worked example: MidState’s route cash confirmations
- How external auditors and examiners read a judgmental selection
- Common mistakes
Three kinds of selection, and what each one supports
The external audit standards draw a line that internal auditors should borrow. Audit sampling, in AICPA AU-C 530 and the PCAOB’s AS 2315, means applying a procedure to less than all of a population in a way that gives every item a chance of selection, so that the results can be projected to the population; the PCAOB’s amended AS 2315 takes effect on 15 December 2026 and keeps the distinction. Selecting specific items, the high-value ones, the unusual ones, the ones a risk assessment points at, is a different procedure that supports conclusions about the items selected and about the risks they represent, but not a projection to the whole. Both are legitimate. The third kind, haphazard or convenience selection, is what happens when the auditor takes items without a defined method, and it supports nothing, because nobody can say what the items represent.
| Kind | How items are chosen | What the results support | When it is the right choice | What a reviewer asks |
|---|---|---|---|---|
| Statistical (random, systematic, monetary-unit) | From a defined, reconciled population, by a method that gives each item, or each dollar, a known chance | A conclusion about the population: an exception rate, a projected misstatement, with a stated confidence | The objective is a population-level conclusion on operating effectiveness or on the amount of error | Population definition and completeness; method; size rationale; evaluation |
| Judgmental, risk-directed | Specific items chosen against explicit criteria: value, risk attributes, timing, source, prior history | A conclusion about the items tested and about the risk the criteria targeted: whether the controls held where the exposure was highest, whether a suspected pattern exists | Concentrated value or risk; small populations; testing design; investigating a known problem; regulatory expectations for specific items | The criteria, the coverage achieved, and whether the conclusion stayed within what the items support |
| Haphazard or convenience | No defined method: what was on top, what the auditee provided, what was easy to obtain | Nothing beyond the items themselves; no basis to say what they represent | Never, for a conclusion; at most for gaining an initial understanding before a real selection | “How were these chosen?” — and the answer ends the discussion |
The word “haphazard” has a technical meaning in the sampling literature, selection without conscious bias that is treated as approximately random, and internal auditors should avoid the term because reviewers read it as “no method”. If a selection was meant to approximate random, use a random method; the mySampler tool draws one in seconds. If it was meant to target risk, say so and say how. The middle ground, where the auditor meant to be random but picked from the top of the file, is the one that fails review.
When judgmental selection is the right call
Judgmental selection is not the fallback for auditors who cannot be bothered with a random sample. In the situations below it is the better method, and a random sample would be the weaker choice, because the objective is not a population-level rate but a question about where the risk sits.
| Situation | Why judgment beats random | Selection criteria that work | Conclusion the selection supports |
|---|---|---|---|
| Value is concentrated | Twenty items are 70 percent of the value; a random sample of 25 from 4,000 would probably miss them all | All items above a threshold, or the top n by value, plus a random sample of the remainder if a population conclusion is also needed | The controls operated over most of the value at risk |
| Small population | Fewer than 30 items: test all, or all the ones that matter | 100 percent, or all items meeting a risk attribute | A conclusion on the whole population, because the whole population was tested |
| Testing design or understanding | One or two transactions walked end to end show whether the control would work; randomness adds nothing | A typical item and an atypical one (an exception, a manual override, a period-end entry) | Design adequacy; not operating effectiveness |
| Investigating a known problem | A prior finding, a complaint, an analytic anomaly: the question is whether the pattern exists, not how often | Items with the anomaly’s attributes: the depot, the user, the date range, the vendor | Whether the pattern is confirmed and its extent within the criteria |
| Regulatory or contractual specific items | Examiners and contracts often specify what must be looked at | The items named: high-risk customers, related-party transactions, the largest exposures | Compliance for the specified items |
| Population cannot be defined reliably | Without a complete population a random sample has no meaning; targeted selection from what exists is honest about the limitation | Items from the available records with the limitation stated | A limited conclusion with the limitation disclosed; and a finding about the missing population |
| Timing risk | Period-end, weekends, holidays, and cutover dates carry more risk than the rest of the year | All items in the risk window | Whether the control held at the point of highest risk |
| Unusual attributes | Round amounts, duplicates, one-time vendors, entries by unusual users, split transactions | Analytics-driven lists of items with the attribute; see the journal entry analytics guide | Whether the attribute indicates a control gap or fraud; extent within the list |
The pattern across the table is that judgmental selection answers “where is the risk and did the control hold there”, while statistical sampling answers “how often does the control fail”. Most engagements need both answers, which is why the strongest designs, discussed under conclusions below, combine a targeted selection of the items that matter most with a random sample of the rest. The risk-directed sampling guide covers the design choices in depth.
The seven elements of a judgmental selection that survives review
A judgmental selection is defensible when a reviewer who was not there can reconstruct why each item was chosen and what the items together represent. That takes seven elements in the workpaper, written before the testing rather than after it, because criteria written after the results are known are criteria fitted to the answer.
| Element | What it records | Why a reviewer needs it |
|---|---|---|
| 1. Objective | What question the selection is meant to answer: design, operation at the highest-risk points, existence of a suspected pattern | The objective determines whether judgment is the right method and what the conclusion may say |
| 2. Population | The full set the items were chosen from, its source, the count and value, and the completeness check performed | Judgmental selection still needs a population; “the ten largest” means nothing without the list they were largest in |
| 3. Selection criteria | The explicit rules: thresholds, attributes, date windows, risk scores, in the order applied | Criteria are what separate risk-directed selection from convenience; they must be reproducible |
| 4. The items and their reason codes | Each item listed with which criterion selected it | A reviewer can see that the criteria were actually applied and nothing was added by hand without a reason |
| 5. Coverage achieved | The items as a share of the population by count and by value; the share of the risk attribute covered | “Twenty items covering 71 percent of value” is a statement of what the test reached; “twenty items” alone is not |
| 6. Independence of selection | Confirmation that the auditor selected from the population, not from items the auditee proposed or pre-assembled | Auditee-selected items are the most common way a judgmental selection becomes worthless |
| 7. Conclusion boundary | A statement, written before testing, of what the results will and will not support | Prevents the finished workpaper from claiming a population conclusion the method cannot give |
Element six deserves its own warning. Process owners are helpful people, and “here are some examples for you” is the sentence that has undone more judgmental selections than any other. The auditor obtains the population, applies the criteria, and tells the auditee which items are needed; the auditee retrieves them. Where the population itself comes from the auditee, as it usually does, the completeness check in element two is what makes the selection independent, and the IPE guide covers how to perform it. The workpaper example shows the seven elements laid out on a page.
Language: deliberate selection versus grabbing
The same twenty items can be described in a way that reads as method or in a way that reads as convenience, and reviewers respond to the words. The table pairs the phrasing that fails review with the phrasing that passes; the difference in every row is that the passing version names the population, the criteria, and the coverage.
| Reads as grabbing | Reads as method |
|---|---|
| A sample of 20 invoices was selected for testing. | From the 4,212 invoices posted in the period (AP register, reconciled to the GL, workpaper C-2), the 20 highest-value invoices from vendors created in the period were selected, covering $1.84 million or 71 percent of new-vendor spend. |
| Judgmentally selected 10 journal entries. | All 10 manual journal entries above $100,000 posted in the last three business days of the quarter were selected (population: 1,140 manual entries in the quarter, JE analytics workpaper D-1); these represent 58 percent of period-end manual entry value. |
| Reviewed a selection of user accounts. | All 14 accounts holding the ERP system administrator role and all 9 accounts with vendor-master and payment-release access were selected from the role export dated 30 June (population 612 active accounts); together they are every account able to change or release a payment without a second person. |
| Tested several reconciliations. | The bank reconciliations for the two accounts with the highest transaction volume and the one account that showed unreconciled items in the prior audit were selected for all three months of the quarter (9 reconciliations of 27 performed). |
| Obtained examples of contracts from the procurement team. | From the contract register (247 contracts active at 30 June, reconciled to the legal department’s list), the auditor selected the 12 contracts above $500,000 in annual value and the 5 contracts with related-party counterparties, covering 64 percent of contracted spend. |
| A number of depots were visited. | Four of twelve depots were selected for site visits: the two with the highest override rates in the full-population analysis and the two acquired depots not yet on the ERP; together they hold 41 percent of route cash volume. |
Three habits produce the right-hand column. Always state the population with its source and reconciliation reference before describing the selection. Always use a number for the criterion, never a word like “large” or “significant”. And always finish with the coverage, so that the reader knows what share of the thing at risk the test reached. The workpaper best practices guide treats this as a documentation standard rather than a style preference, which is what it is.
What you can and cannot conclude
A judgmental selection supports three kinds of conclusion and cannot support two others, and the workpaper’s conclusion sentence must stay on the right side of the line. It can conclude on the items tested: these twenty invoices were properly approved, or three were not. It can conclude on the risk the criteria targeted: the controls operated over the highest-value new-vendor spend, or they did not at the point of highest exposure. And it can conclude on design, where the objective was design. It cannot state an exception rate for the population, because the items were not chosen in a way that represents it, and it cannot project a misstatement, for the same reason. “3 of 20 exceptions, 15 percent” is a fact about the twenty; “15 percent of invoices are unapproved” is a claim the method does not support, and reviewers know the difference.
| Conclusion type | Supported by judgmental selection? | Example wording |
|---|---|---|
| On the items tested | Yes | “All 20 selected invoices were approved by an authorized approver before payment.” |
| On the targeted risk | Yes, within the criteria | “Approval controls operated over the 71 percent of new-vendor spend represented by the selection; the control was not tested for lower-value invoices.” |
| On design | Yes, where that was the objective | “The walkthrough of two transactions confirmed the control is designed to require a second approver above $10,000.” |
| Population exception rate | No | Not: “15 percent of invoices lacked approval.” Instead: “3 of the 20 highest-value new-vendor invoices lacked approval; the extent across the remaining population was not tested and is addressed by the random sample at C-4.” |
| Projected misstatement | No | Not: “unapproved spend is estimated at $410,000.” Instead: “unapproved spend within the selection totaled $61,200; no projection is made.” |
| Operating effectiveness of the control across the period | Only when combined with a statistical sample of the remainder, or when the selection is the full population | “The 20 targeted items and the random sample of 40 from the remaining population together support the conclusion that the control operated effectively; see C-4 for the sample evaluation.” |
The combined design in the last row is the professional answer to most engagements. Take the items that matter most by value or risk in full, then draw a random sample from what remains, evaluate the two parts separately, and conclude on the population from the sample and on the risk from the targeted items. Monetary-unit sampling does something similar automatically, since large items are almost certain to be selected, which is why the sample size guide recommends it for value-weighted populations. What the combined design never does is mix the exceptions from the two parts into one rate; the three targeted exceptions and the one random exception are reported as what they are.
Six rationale paragraphs to copy
Each paragraph below is a complete selection rationale for a workpaper, with the seven elements in order. Replace the bracketed items, keep the structure, and write it before the testing starts.
1. High-value coverage. Objective: to test whether the approval control operated over the invoices carrying most of the value at risk. Population: [n] invoices totaling [$] posted between [dates], per the AP register at [ref], reconciled to the GL at [ref] with no unreconciled differences. Criteria: all invoices above [$threshold], applied in descending value order. Items: [n] invoices listed at [ref] with the criterion recorded against each. Coverage: [n] of [n] invoices, [percent] of value. Independence: the auditor extracted the register and applied the criteria; the auditee supplied documents for the listed items only. Conclusion boundary: results support a conclusion on the selected items and on the control’s operation over [percent] of value; no exception rate or projection to the population is made; population-level conclusions rest on the random sample at [ref].
2. New or changed master data. Objective: to test whether vendor onboarding controls operated for vendors created or changed in the period, where the fraud and error risk concentrates. Population: [n] vendor master records created or with bank details changed between [dates], per the change log at [ref], reconciled to the vendor master count. Criteria: all records with a bank account change, plus all new vendors with payments above [$] in the period. Items: [n] at [ref]. Coverage: 100 percent of bank changes; [percent] of new-vendor spend. Independence: extracted by the auditor from the system. Conclusion boundary: supports a conclusion on onboarding and change controls for the items and the targeted risk; not an exception rate for all vendor changes.
3. Period-end manual journals. Objective: to test whether manual journal entries at period end, the point of highest misstatement risk, were reviewed and supported. Population: [n] manual entries in the quarter per the JE extract at [ref], reconciled to the GL posting total. Criteria: entries above [$] posted in the last [n] business days, entries by users outside the accounting team, and entries with round amounts above [$]. Items: [n] at [ref], each with the criterion. Coverage: [percent] of period-end manual entry value. Independence: extracted and filtered by the auditor. Conclusion boundary: supports conclusions on the items and on the period-end risk; the operating effectiveness of JE review across the quarter is addressed by the sample at [ref].
4. Privileged access. Objective: to test whether the accounts able to bypass segregation of duties are restricted and reviewed. Population: [n] active accounts in [system] per the role export dated [date] at [ref], reconciled to the HR active headcount with [n] differences investigated at [ref]. Criteria: all accounts holding [administrator role] and all accounts holding both [conflicting role A] and [conflicting role B]. Items: [n] accounts at [ref]. Coverage: 100 percent of the privileged population. Independence: extracted by the auditor with IT observing. Conclusion boundary: supports a conclusion on the whole privileged population, since all of it was tested; supports no conclusion about the appropriateness of standard-user access, which is addressed at [ref].
5. Follow-on from an exception. Objective: to determine whether the exception found at [ref] is isolated or indicates a pattern. Population: [n] items sharing the exception’s attributes ([depot, user, vendor, date range]) per [ref]. Criteria: all items in the population, or where the population exceeds [n], the [n] most recent and the [n] highest-value. Items: [n] at [ref]. Coverage: [percent] of the attribute population. Independence: defined by the auditor from the original exception. Conclusion boundary: supports a conclusion on whether the pattern exists within the attribute population and its extent there; supports no conclusion about items outside the attributes.
6. Small population tested in full. Objective: to conclude on the operating effectiveness of [control] for the period. Population: [n] instances of the control in the period per [ref], reconciled to [source]. Criteria: none; the population is below [30] items and was tested in full. Items: all [n] at [ref]. Coverage: 100 percent. Independence: population obtained by the auditor from [system]. Conclusion boundary: supports a conclusion on the population, since the population was tested; any exception is an exception in the population, not a sample result.
Worked example: MidState’s route cash confirmations
MidState Beverage’s FY27-01 route cash audit used both kinds of selection deliberately, and the workpapers say which was which. For the settlement reconciliations, the objective was a population-level conclusion on whether the daily control operated, so the team drew a stratified random sample: 60 reconciliations, five per depot, from the year’s 37,400 settlements, evaluated as a sample; 14 failed, and the conclusion on operating effectiveness rested on that sample and the full-population override analysis. For customer balances, the objective was different: the FY26 diversion had run for five months because the affected customers received no statements and never noticed, so the question was whether a similar loss could be sitting undetected now, at the customers where it would most likely sit. That is a judgmental question, and the selection rationale in the workpaper read as follows.
Objective: to determine whether cash-paying customers who receive no monthly statement carry balances that differ from the ledger, indicating diverted or misapplied payments. Population: 1,130 customers flagged in the customer master as receiving no statement (workpaper F-2, reconciled to the 4,200 cash-paying customers at F-1), with ledger balances totaling $9,421,880 at 30 April. Criteria: the 12 customers with the largest ledger balances, plus the 8 customers with the highest cash volume at the four depots with the highest override rates in the full-population analysis (D-3). Items: 20 customers listed at F-3 with the criterion against each. Coverage: 20 of 1,130 unstated customers; 38 percent of the unstated ledger balance; four depots holding 41 percent of route cash volume. Independence: the auditor selected from the customer master extract; depot staff were not told which customers would be confirmed until the confirmations were sent by the audit team directly. Conclusion boundary: results support a conclusion on the 20 customers and on the targeted risk, whether large unstated balances at high-override depots are misstated; they support no exception rate for the 1,130, and the population-level question of statement coverage is a finding in its own right (F5) regardless of confirmation results. A monetary-unit sample of 70 deposits (workpaper E-1) separately supports the conclusion on deposit completeness across the population.
The rationale did its job in two directions. When the confirmations came back, the team could say exactly what the results meant: the 20 balances agreed within tolerance, which showed no diversion at the highest-risk points and did not show that the other 1,110 customers were fine; the finding that 1,130 customers received no statement stood on its own as a control gap. And when the COO argued at the closing meeting that the clean confirmations proved the statement gap was harmless, the conclusion boundary, written weeks earlier, was the answer: the confirmations were selected to find a loss where it was likeliest, not to measure the population, and a control that would let a loss run undetected at any of 1,130 customers is a control gap whether or not a loss was running that month. The root cause guide covers how the finding was written; the substantive testing guide covers the confirmation procedure itself.
How external auditors and examiners read a judgmental selection
Two outside readers judge internal audit’s selections, and both apply the distinction this guide rests on. An external auditor deciding whether to use internal audit’s work under PCAOB AS 2605 or the AICPA equivalent asks whether the testing supports the conclusion drawn from it; a judgmental selection presented as an operating-effectiveness conclusion will be discarded and re-performed, while the same selection documented with its criteria, coverage, and boundary, alongside a random sample of the remainder, is work the external auditor can use. Bank and insurance examiners read files the same way, and an examiner who finds “management provided examples” in a selection rationale reads it as a finding about the audit function rather than about the process. The regulator expectations guide covers what examiners look for in workpapers; the short version is that they look for exactly the seven elements above.
There is one more reader: the function’s own quality program. A self-assessment or external quality assessment reviews engagement files against the Standards’ evidence requirements, and undocumented selection is among the most common file-level observations assessors raise, because it is visible on every page and easy to fix. A function that adopts the seven elements as a documentation standard, checks them at the fieldwork QC checkpoint described in the QAIP kit, and puts the rationale in the sampling memo removes the observation before it is made.
Common mistakes
| Mistake | What it looks like | Fix |
|---|---|---|
| “Judgmental” as a synonym for undocumented | The word appears in the workpaper with no criteria | Criteria, coverage, and boundary written before testing |
| Auditee-provided examples | “Management provided ten contracts for review” | The auditor obtains the population and selects; the auditee retrieves |
| Criteria written after the results | Selection rationale added at review time to fit what was tested | Rationale dated before fieldwork; the reviewer checks the date |
| Rates from targeted items | “3 of 20, a 15 percent exception rate” | Report the exceptions as facts about the items; rates come from samples |
| No population | “The largest items were selected” with no list they were largest in | Population, source, count, value, completeness check, every time |
| Coverage unstated | Twenty items, and no idea what share of value | Count and value coverage in the rationale and the conclusion |
| Judgment where a rate was needed | A SOX operating-effectiveness conclusion drawn from the ten largest transactions | Random sample of the population, or targeted items plus a sample of the remainder |
| Random where judgment was needed | A random 25 from 4,000 invoices that misses every large one, and a clean conclusion over 4 percent of value | Targeted high-value items first; then the sample |
| Mixing the two sets of exceptions | Targeted and sampled exceptions combined into one rate | Evaluate separately; report separately |
| The word “haphazard” | Used to mean “we picked some” | Use a random method or a documented judgmental one; never the word |
Judgmental selection is the auditor’s most powerful tool for putting testing where the risk is, and the least respected, because it is so rarely written down properly. Seven elements, written before the work, and a conclusion that stays within what the items support: that is the whole discipline, and a selection documented that way survives any reviewer, because the reviewer can see exactly what was done and exactly what it means.
Related guides
- The sampling memo template
- Random sampling in Excel and Google Sheets
- Audit sample sizes: 25, 40, 60
- Risk-directed sampling techniques
- mySampler (free tool)
- Audit evidence
- Testing information produced by the entity
- Audit workpaper example
- Workpaper best practices
- Substantive testing for beginners
- Journal entry testing
- Journal entry analytics
- Root cause analysis for audit findings
- Start here
Leave a Reply