Here is the answer up front, because you have probably asked the question and been handed “it depends”: the famous audit sample sizes are rounded outputs of attribute-sampling math run at zero expected deviations. 25 comes from 90% confidence that the deviation rate does not exceed 10% (the raw math says 22). 40 comes from 95% confidence at a 7.5% tolerable rate (raw: 39). 60 comes from 95% confidence at a 5% tolerable rate (raw: 59). That is the whole secret. Everything else — why those parameters, when the conventions apply, and why finding a single exception quietly breaks the plan — is what this guide covers.
This is the most-asked, worst-answered question in internal audit. Juniors inherit the numbers as folklore (“we always do 25”), seniors defend them by citing “methodology,” and almost nobody can produce the derivation when a sharp auditee — or a sharper regulator — asks why 25 items prove anything about a population of 40,000. By the end of this piece you will be able to derive the numbers yourself, know exactly when they mislead, and document a defensible choice for any engagement.
In this guide
- The math in one formula
- The full sample-size grid, computed exactly
- Where 25, 40, and 60 — and the frequency ladder — come from
- The rule nobody tells juniors: one exception breaks the plan
- Five situations where the conventions mislead
- A defensible decision grid for non-SOX audits
- Documenting the decision (and doing it in sixty seconds)
The math in one formula
Attribute sampling answers one question: how confident can I be that the true deviation rate in this population is at or below some tolerable level, given what I saw in my sample? When you plan to accept the control only if you find zero deviations, the sample size collapses into a single line of algebra. If the true deviation rate were exactly the tolerable rate (call it TDR), the probability that a random sample of n items shows no deviations at all is (1 − TDR)ⁿ. You want that probability to be small — no more than (1 − CL), where CL is your confidence level. Solve for n:
n ≥ ln(1 − CL) ÷ ln(1 − TDR)
Walk one through. Say 90% confidence and a 10% tolerable deviation rate: ln(0.10) ÷ ln(0.90) = 21.85, so n = 22. If a control actually failed one time in ten, a random sample of 22 clean items would occur less than 10% of the time — so a clean 22 lets you assert, with 90% confidence, that the failure rate is at or below 10%. Round 22 up to a tidy 25 and you have the most famous number in SOX testing. The rounding is not decoration, either: at n = 25, a clean sample actually delivers about 93% confidence at the 10% tolerable rate. The conventions carry a small built-in cushion.
Notice what is absent from the formula: the population size. For large populations, the binomial math barely cares whether you are sampling from 5,000 items or 5 million — which is why “but the population is huge” is not an argument for a bigger sample, and why small populations (where the formula’s assumption breaks) get their own treatment below.
The full sample-size grid, computed exactly
Here is the zero-expected-deviation grid across the parameter pairs that matter in practice. These are exact binomial results from the formula above (we computed them fresh for this article — no table was harmed by transcription error):
| Tolerable deviation rate | Sample size at 90% confidence | Sample size at 95% confidence |
|---|---|---|
| 2% | 114 | 149 |
| 3% | 76 | 99 |
| 4% | 57 | 74 |
| 5% | 45 | 59 (→ the “60”) |
| 6% | 38 | 49 |
| 7.5% | 30 | 39 (→ the “40”) |
| 8% | 28 | 36 |
| 10% | 22 (→ the “25”) | 29 (→ the “30”) |
| 15% | 15 | 19 |
| 20% | 11 | 14 |
Read the grid vertically and the economics of assurance jump out: tightening the tolerable rate is expensive (5% costs roughly triple what 10% costs), while stepping confidence from 90% to 95% costs only about a third more items. That asymmetry is worth remembering when someone reflexively demands “more confidence” — the tolerable rate, not the confidence level, is usually what is doing the real work in your conclusion.
Where 25, 40, and 60 — and the frequency ladder — come from
The specific pairings became convention through external audit’s ICFR methodologies in the post-SOX era and migrated into internal audit wholesale. The profession’s reference texts — the AICPA’s Audit Sampling guide and the sampling standards for external auditors (PCAOB AS 2315, which, for the standards-watchers, was amended with an updated version effective December 15, 2026) — supply the statistical machinery; the firms’ testing matrices supplied the rounded numbers. What emerged is the frequency ladder every SOX shop recognizes:
| Control frequency | Population per year (approx.) | Conventional sample | What it implicitly assumes |
|---|---|---|---|
| Annual | 1 | 1 | Not sampling — you are testing the control’s single operation end to end |
| Quarterly | 4 | 2 | Coverage-based selection, not statistics; the binomial formula does not apply this small |
| Monthly | 12 | 2–5 | Same — judgmental coverage of the year |
| Weekly | 52 | 5–15 | Transitional zone; some shops apply a small-population adjustment |
| Daily | ~250 | 20–40 | Approaching the binomial zone; 25 at moderate assurance is the workhorse |
| Multiple times per day | Thousands+ | 25 / 40 / 60 | Full attribute-sampling territory: 90/10, 95/7.5, or 95/5 at zero deviations |
Understand what the ladder actually is: the bottom rows are statistics; the top rows are coverage judgment dressed in a table. Testing 2 of 4 quarterly reconciliations gives you no statistical confidence about anything — it gives you direct evidence about half the population, which is a different (and for small populations, better) kind of assurance. The ladder works because frequency is a proxy for population size, and population size determines which regime you are in. Trouble starts when people quote “we tested 25” reverently for a population of 30, or treat 2-of-4 as if it projected to a rate.
The rule nobody tells juniors: one exception breaks the plan
Every size in the grid above is a zero-deviation plan: the sample only supports the stated conclusion if it comes back clean. Find one exception in your 25 and you cannot shrug, note “96% pass rate,” and conclude the control is effective at the original parameters. The math is unforgiving: to conclude at 90% confidence and a 10% tolerable rate while accepting one observed deviation, you needed a sample of 38, planned that way from the start. At 95% and 5%, tolerating a single deviation moves the requirement from 59 items to 93.
So when an exception surfaces, you have exactly three honest moves:
- Investigate before anything else. Is it a deviation at all, or a documentation artifact? If real, is it isolated (one preparer, one system migration week, one entity) or does the cause imply a systematic failure? Root cause decides everything downstream — an exception whose cause touches the whole population cannot be averaged away by any sample size.
- Expand rationally, not ritually. Extending to the one-deviation sample size (38, or 93, per the parameters) is legitimate only if the investigation supports an isolated-error story and the expansion is selected randomly from the same population. “Test five more and hope” is neither statistics nor judgment — it is theater, and experienced reviewers read it as exactly that.
- Conclude the control failed the test. Often the right answer. A real, unexplained deviation in a small zero-deviation sample is evidence the deviation rate may exceed tolerance — report it, rate it, and let issue validation confirm the fix later. The sample did its job: it found the problem.
If your methodology allows “expected deviations greater than zero” plans up front — sensible for controls with known noise — the grid shifts accordingly (and the sizes grow fast). The discipline that matters is deciding the acceptance number before selecting the sample, not after seeing what came back.
Five situations where the conventions mislead
1. Small populations. The binomial formula assumes each draw barely dents the population. For a control that operates 24 times a year, “sample 25” is a logical impossibility and “sample 22” is absurd — you are in coverage territory: test a meaningful fraction (or everything) and say so. Between roughly 50 and 250 items, a finite-population correction meaningfully shrinks the required sample; our free mySampler applies the AICPA finite-population adjustment automatically when you give it the population size.
2. Non-homogeneous populations. The math assumes one population with one deviation rate. If wires above $1 million follow a different approval path than wires below it, that is two populations wearing one label — sample them separately or stratify. Mixing strata and quoting one rate is the most common silent invalidator of otherwise tidy workpapers.
3. Non-random selection. The confidence statement is a property of random selection. Haphazard picking (“grab some from each month”) feels random and is not — humans oversample the legible, the recent, and the middle of lists. If you claim statistical confidence, the selection must be reproducibly random: a seeded random number generator, with the seed documented so a reviewer can regenerate your exact sample. (This is precisely why mySampler stamps a seed on every selection it exports.) If you selected judgmentally — targeting high-risk items on purpose — that can be excellent auditing, but write your conclusion as judgmental coverage, never as a projected rate.
4. Rates when you needed dollars. Attribute sampling speaks about how often a control fails, not how many dollars are wrong. If the engagement question is monetary — “is this account materially misstated?” — you want monetary-unit sampling or another variables technique, where selection probability follows dollar value. Testing 25 invoices as attributes and opining on the account balance is a category error that survives in far too many files.
5. A thoughtless tolerable rate. The 10% behind the classic 25 means: one failure in ten is acceptable. Say that sentence out loud about the control you are testing. For a routine expense-report approval, maybe. For wire-transfer callbacks or privileged-access grants, a 10% failure tolerance is indefensible — which pushes you down the grid toward 5% (n = 59) or tighter, or toward testing the full population with analytics instead of sampling at all. The tolerable rate is a risk judgment wearing a percentage; own it consciously.
A defensible decision grid for non-SOX audits
Internal audit outside SOX has a freedom external auditors lack: the sample serves your engagement objective, not an ICFR opinion template. That freedom is only defensible if you can articulate the objective-to-parameters logic. This is the grid I use, and would happily defend to a quality assessor:
| Assurance objective | Suggested parameters | Sample (zero-deviation plan) | Notes |
|---|---|---|---|
| High reliance on a critical control (wires, privileged access, safety-critical steps) | 95% confidence, 5% tolerable | 59–60 | Consider 100% testing via analytics instead — populations this critical are often small enough or data-rich enough |
| Standard reliance on a key control | 90% confidence, 10% tolerable | 22–25 | The workhorse; the classic 25 with its cushion |
| Moderate comfort on a secondary control | 90% confidence, 15% tolerable | 15 | Honest label for what many “sample of 15” tests already are |
| Directional read for an advisory or scoping exercise | 85–90% confidence, 20% tolerable | 9–11 | Fine for temperature-taking; never dress it up as assurance |
| Zero-tolerance condition (fraud indicators, sanctions hits) | Discovery sampling: size to detect a rate of 1–2% at 90–95% | ~115–300 | Any single hit triggers investigation; sampling here is a detection net, not a rate estimate |
| Population under ~50 | Coverage judgment | 25–100% of items | State coverage achieved (“tested 6 of 12 months, both preparers”); make no rate projection |
Two habits make this grid stick in a real department. First, put the mapping in your methodology once — “our standard parameters by reliance level are…” — so engagement teams inherit consistency instead of renegotiating statistics every fieldwork. Second, when risk is concentrated, split your approach explicitly: a judgmental slice of the scary items (documented as judgmental) plus a random sample of the remainder (documented with parameters). Our deeper guides on the foundations of random sampling and sampling strategies for high-risk areas unpack both habits at length.
Documenting the decision (and doing it in sixty seconds)
A sample that cannot be reperformed is an anecdote. The workpaper standard worth holding your team to — for every sample, statistical or judgmental — is five lines: the population and how you assured its completeness; the method (attribute, MUS, discovery, judgmental) and why it fits the objective; the parameters (confidence, tolerable rate, expected deviations) or the judgmental criteria; the selection mechanism with its seed; and the evaluation rule decided in advance, including what happens on an exception. Five lines, and any reviewer — internal QA, external assessor, regulator — can regenerate your sample and re-trace your logic.
This is exactly the workflow our free mySampler tool automates: paste a population or range, pick the method, set attribute parameters (with the finite-population correction applied when you provide the population size), select with a seeded generator, and export a sampling memo carrying the method, parameters, seed, and timestamp — workpaper-ready, reperformable, and free, with everything running in your browser. And if spreadsheets are your habitat, our guide to random sampling in Excel and Google Sheets shows the reproducible-selection mechanics by hand.
Final Thoughts
The numbers were never magic. 25, 40, and 60 are rounded solutions to one line of algebra — 90/10, 95/7.5, and 95/5 at zero expected deviations — conventions that earn their keep only when their assumptions hold: a large, homogeneous population, genuinely random selection, and a tolerable rate someone actually chose. Learn the formula, keep the grid handy, refuse to average away exceptions, and label judgmental work as judgment. Do that, and “why 25?” stops being the question you dread from auditees and becomes the one you hope they ask.
Primary references: the AICPA Audit Sampling guide (the profession’s standard tables and the finite-population correction) and PCAOB AS 2315, Audit Sampling. All computed values in this article are exact binomial results from the stated formula.
Leave a Reply