,

Management Review Controls: Documenting and Testing the Hardest Controls in SOX

Management review controls are where SOX programs lose the most arguments with their external auditors, and the arguments are almost always the same one: a control that management believes is effective because a competent person looked at the numbers, and an auditor who cannot see what the person looked at, what they expected, what would have caught their eye, or what they did about it. The control may be excellent. The evidence that it is excellent usually does not exist, because nobody designed the review to leave evidence, and the auditor is left with a signature and a conversation. This guide sets out how to design a review control that can be tested, how precision is established and documented, what evidence of review has to look like, how to test a review control on design and operation, and how review-control deficiencies are evaluated, with a documentation template and a worked example at a bank.

This guide was rewritten in September 2026 from a shorter version published in August 2026. Its framework is the PCAOB’s Staff Audit Practice Alert No. 11 of October 2013, which set out the precision factors auditors apply to review controls and remains the reference inspectors and engagement teams quote, together with AS 2201 and the SEC’s 2007 interpretive guidance for management.

In this guide

Why review controls are the hardest controls in the program

Most key controls are either mechanical, so the evidence is the mechanism, or transactional, so the evidence is the transaction. A review control is a judgment applied to information, and judgment leaves no trace unless the design forces it to. That is the first difficulty. The second is that review controls are asked to do the most work in the matrix: they are the detective layer over estimates, over the close, over entity-level risks, and over every process where the preventive controls are known to be imperfect, which means they are frequently cited as compensating controls in deficiency evaluations and as the controls that address fraud risk and management override. The third is regulatory history. The PCAOB’s inspection findings on audits of internal control in the years after 2010 concentrated on two things: auditors who had not tested the completeness and accuracy of the information a review used, and auditors who had not evaluated whether the review operated at a level of precision that would catch a material misstatement. Staff Audit Practice Alert No. 11 of October 2013 wrote both into a list of factors, and external auditors have tested against that list ever since. Management programs that were built before the alert, and many built after it, describe review controls as “the controller reviews the reconciliation and signs it”, which answers none of the factors.

The result is a predictable annual cycle: the auditor asks for the reviewer’s evidence, management produces the sign-off, the auditor asks what the reviewer would have caught, management explains the reviewer’s experience, and the control is either downgraded, re-tested by the auditor at greater cost, or written up as a deficiency with a remediation plan that says “enhance documentation”. The way out is to design the review so that its precision is a property of the control rather than of the person, and so that its performance produces evidence as a by-product. Everything below is in service of that.

The anatomy of a review control

Every review control that works has six parts, and a review that is missing any of them is a look rather than a control. The parts are the design questions the control description has to answer, and the table pairs each with the question and with what a weak answer sounds like.

ElementDesign questionStrong answerWeak answer
Information reviewedWhat exactly does the reviewer look at, produced by whom, from what source, at what level of detail?Named report or analysis with source, parameters and level of aggregation (by product, location, account)“The financial results”
ExpectationWhat does the reviewer expect to see, and where does the expectation come from?Budget, forecast, prior period adjusted for known changes, independent data (volumes, rates), an analytical modelThe reviewer’s experience
Comparison and thresholdWhat comparison is made, and what difference triggers investigation?Stated threshold in amount and percentage, set below the level at which a misstatement would matter for the assertion“Significant” or “unusual” variances
InvestigationWhat does the reviewer do about an item over the threshold?Obtains explanation supported by evidence, corroborates independently of the preparer, documents the conclusionAsks the preparer, accepts the answer
ResolutionWhat happens when the investigation finds a problem?Correction posted, or a documented decision not to correct with the reason and the amount; escalation path definedNot defined
EvidenceWhat record shows each of the above happened?Annotated analysis, review notes or workflow record showing items questioned, explanations, corroboration and conclusions; sign-off with dateSignature and date

The expectation element is the one most often missing and the one that matters most, because a review without an independent expectation can only detect what looks wrong to the reviewer, and misstatements do not announce themselves. A variance analysis against budget is a review with an expectation; a read-through of the trial balance is not, however senior the reader. Where a review has no external expectation available, the design has to supply one, through a model, a ratio the reviewer computes, or a reconciliation to an independent source, or the control has to be described as what it is, a reasonableness check, and rated accordingly.

Precision: the six factors and how to document them

Precision is the property that decides whether a review control can be a key control at all: would this review, performed as designed, prevent or detect a misstatement that could be material to the assertion it addresses? Staff Audit Practice Alert No. 11 listed the factors that determine the answer, and the program should document each factor for every key review control, because they are the questions the external auditor will ask in that order.

Factor (Alert 11)More preciseLess preciseHow to document and test
Objective of the reviewDesigned to prevent or detect misstatements in a named assertionDesigned to identify and explain differences for management informationState the assertion in the description; test that investigations are framed as error detection, not business commentary
Level of aggregationBy product, location, account or transaction type, at a level where a misstatement would stand outTotal company, total revenue, consolidated marginShow the aggregation level in the information reviewed; compute whether a misstatement of the tolerable size would move the reviewed figure by more than the threshold
Consistency of performancePerformed every period, same procedure, same reviewer role, same thresholdPerformed when time allows, procedure varies, threshold variesEvidence of each performance in the period; compare procedures across instances
Correlation to the assertionDirectly addresses the assertion (a revenue cut-off review looks at cut-off)Indirectly related (a margin review is asked to detect cut-off)Map the review to the assertion in the matrix; test that the investigation would surface the specific misstatement type
Predictability of expectationsExpectation from an independent, reliable source; relationships stable enough that deviations are meaningfulExpectation is the prior period or the preparer’s own number; relationships volatileRecord the expectation’s source; test its independence from the information reviewed and its reliability
Criteria for investigationThreshold well below the amount that would be material to the assertion; both amount and percentage; applied to each lineThreshold near materiality; percentage only on large bases; applied to totalsCompare the threshold to tolerable misstatement for the assertion; test that items over threshold were investigated

The arithmetic behind the aggregation and threshold factors is worth doing explicitly, because it converts an argument into a calculation. If interest income is reviewed at the total-bank level and the review threshold is 3 percent, the review would investigate a variance of roughly $9.5 million on annual interest income of $315 million; a $4 million misstatement, below Lakeshore Bancorp’s overall materiality of $6.5 million but well inside the range that matters once deficiencies are combined, would be invisible. The same review at portfolio level, with a threshold of $250,000 per portfolio and 2 percent, would surface a $4 million error in any single portfolio and most errors above $500,000. The documentation records the calculation, the threshold chosen, and the conclusion that the review is precise enough to serve as a key control for the assertions listed, or is not and serves as a monitoring control instead. The SOX scoping guide sets out where the threshold numbers come from, and the deficiency evaluation guide uses the same precision test when a review is offered as a compensating control.

Frequency belongs in the precision analysis as well, although Alert 11 does not list it separately. A quarterly review over an account that moves daily detects errors up to a quarter late, which is acceptable for an annual assessment only if the errors would still be visible at quarter-end rather than reversed, netted or buried by then; a monthly review over the same account detects them within the month and before the quarterly financial statements. The design record states why the frequency chosen is sufficient for the assertion, and the answer is usually about when the misstatement would become invisible, not about the reviewer’s calendar.

Writing the control description so it can be tested

A testable description names all six elements of the anatomy and states the precision factors in numbers. It reads, in full, something like: “Within eight business days of month-end, the Controller reviews the interest income variance analysis (report FIN-214, generated from the general ledger by product portfolio, compared to the budget and to the prior month adjusted for average balance and rate changes from the ALM system). Variances greater than $250,000 or 2 percent at portfolio level are investigated by obtaining the preparer’s explanation, corroborating it to the loan system’s average balance and yield data, and recording the explanation and conclusion in the review workbook. Errors identified are corrected by journal entry before the close is finalized or documented as uncorrected with the amount. The Controller signs the workbook and the Chief Accounting Officer reviews the workbook and the list of uncorrected items quarterly.” Every clause in that description is an attribute the tester can inspect, and every number is one the reviewer can be held to. The control descriptions guide works through the drafting rules, including the one that most often fails here, that the description must name the report and its parameters rather than “the analysis”.

The information inside the review

A review control is an IT-dependent manual control by construction: the reviewer applies judgment to information a system, a spreadsheet or a preparer produced, and the review is only as good as the completeness and accuracy of that information. Alert 11 made the point directly, describing audits that tested the reviewer and never tested the report, and the program treats every report and workbook inside a key review as information produced by the entity, tested once and cross-referenced. Three questions apply to each: is the information complete (every portfolio, every account, every period), which is proved by reconciling the analysis to the ledger; is it accurate, which is proved by tracing sampled lines to source and re-performing calculations; and were the parameters right on the instance the reviewer used, which is proved by inspecting that instance, not a fresh run. Where the analysis is prepared in a spreadsheet, the spreadsheet’s logic and change control are part of the control, per the end-user computing guide. The IPE testing guide carries the full method and the control classification guide explains why the two-part test cannot be collapsed into one.

The expectation also has an information dependency that is easy to miss. A budget is only a useful expectation if the budget itself was prepared independently of the numbers it is being compared against and is not revised to match actuals; a “flexed” budget that moves with volume is a better expectation than a static one but only if the volume data is independent of the ledger. The tester asks where the expectation came from, and for a review whose expectation is the prior period, asks how the reviewer would notice an error that has been in the number for two periods.

Evidence of review: what has to exist

The standard for evidence of review is that a tester, reading the file without talking to the reviewer, can see which items exceeded the threshold, what the reviewer asked about each, what answer was obtained, how it was corroborated, what the reviewer concluded and what was corrected. A signature attests that the review happened; it does not evidence any of those things, and an auditor who accepts it is relying on inquiry, which the standards say is not sufficient on its own for a control of this kind. Four forms of evidence meet the standard. Annotated analyses, where the reviewer’s questions, the explanations and the conclusions are written on or alongside the flagged lines. Review workbooks or logs, where each item over threshold is a row with explanation, corroboration reference, conclusion and correction. Workflow records, where the review runs in a close management or reconciliation tool that captures comments, attachments and approvals per item. And meeting records, for reviews performed in a committee, where the minutes record the items discussed and the decisions, supported by the pack presented.

Two kinds of evidence do not meet the standard and appear constantly. Email chains between the reviewer and the preparer show that questions were asked but rarely show corroboration or conclusion, and they are not retained systematically; they can support a review log but cannot replace it. And “reviewer’s notes” that are prepared by the preparer, listing explanations for every variance before the reviewer has seen them, show the preparer’s diligence, not the reviewer’s; they are common, and the tester should ask who wrote the explanation and when. A related trap is the review that identifies items and never records that any were investigated because “there were none over threshold”; the evidence in that case is the thresholded analysis itself, showing that no item exceeded the criteria, which is fine if the analysis is retained and the threshold is visible.

Timeliness is evidence too. A review performed after the close is final detects errors the financial statements already contain; the description states the deadline relative to the close and the evidence carries dates that meet it. And where the reviewer is also the person who could be overriding controls, the review needs a second reviewer or an independent view of the uncorrected items, which is the management override consideration the SOX 404 guide discusses under entity-level controls.

Testing a review control: design and operation

The design test is a walkthrough with the reviewer, not the preparer, on a real instance. The tester watches the reviewer perform the review or re-performs it alongside them, and evaluates the six anatomy elements and the six precision factors against what the reviewer actually does, which often differs from the description in both directions: reviewers frequently apply a tighter threshold than documented, and frequently skip the corroboration step the description promises. The design conclusion states whether the control, as performed, would detect a misstatement of the tolerable size in the assertion, and it names the reviewer’s competence and authority as part of the conclusion, because a review performed by someone who could not recognize the error or could not force the correction is deficient in design.

The operating test samples instances across the period at the frequency-based sizes (two quarterly, two to five monthly, per the sample sizes guide) and, for each instance, tests every element: the information reviewed was the right report with the right parameters and reconciles to the ledger; the expectation was the documented source; items over threshold were identified completely (the tester recomputes which lines exceeded the threshold); each was investigated with a documented explanation; the explanation was corroborated to something other than the preparer’s word (the tester re-performs the corroboration on selected items); conclusions were recorded; corrections were posted or uncorrected items documented; the review met the deadline; and the sign-offs are by the named roles. The re-performance step is what distinguishes a test of the review from a test of the file: for a selection of the investigated items, the tester obtains the same supporting data and reaches their own conclusion, and a difference between the tester’s conclusion and the reviewer’s is an exception whether or not the reviewer’s conclusion turned out to be right.

The test program

#AttributeMethodException if
1Information reviewedInspect the instance reviewed; confirm report identity, parameters and level of aggregation; reconcile to the ledgerWrong report, parameters or period; analysis does not reconcile; aggregation higher than described
2ExpectationTrace the comparison basis to its documented source; assess independence from the reviewed figuresExpectation is the reviewed number restated, or was revised to match actuals
3Threshold applied completelyRecompute items exceeding the threshold on the instance; compare to items the reviewer flaggedAny item over threshold not identified
4Investigation documentedInspect the explanation recorded for each flagged itemMissing or generic explanation (“timing”, “volume”) with no specifics
5CorroborationInspect evidence that the explanation was verified independently of the preparer; re-perform on selected itemsExplanation accepted on the preparer’s word; re-performance disagrees
6Conclusion and resolutionInspect the conclusion per item; trace corrections to journal entries; inspect the uncorrected items listNo conclusion; identified error not corrected or documented
7TimelinessCompare review date to the deadline in the description and to the close finalization dateReview after the close was finalized
8Performer and reviewerConfirm the roles who performed and approved match the description; assess competence and authorityPerformed by a delegate outside the description; reviewer lacked authority to require correction
9Precision (design)Document the six factors; compute the threshold against tolerable misstatement at the aggregation level usedReview could not detect a misstatement of the tolerable size in the assertion
10ConsistencyCompare procedure, threshold and evidence across sampled instancesProcedure or threshold varies between instances without a documented change

Review-control deficiencies and how they are evaluated

Review controls generate a characteristic set of deficiencies, and each maps to a classification question. A precision deficiency, where the review is performed well but could not catch a misstatement of the size that matters, is a design deficiency over the assertion, and its severity is driven by whether any other control addresses the assertion at adequate precision; where the review was the only control, the plausible misstatement is bounded by the review’s threshold at its aggregation level, which is often material. An evidence deficiency, where the review is performed but the file does not show it, is treated by most auditors as an operating deficiency because the control cannot be shown to have operated; its severity depends on whether the tester’s re-performance finds errors the review missed. An investigation deficiency, where flagged items were explained but not corroborated, is an operating deficiency whose magnitude is the sum of the items accepted on the preparer’s word. And an information deficiency, where the analysis reviewed was incomplete or wrong, is evaluated as an IPE failure over every control that used the report. The deficiency evaluation guide walks the likelihood and magnitude analysis, and the point specific to reviews is that “enhance documentation” is a remediation for the evidence deficiency only; a precision deficiency is remediated by changing the review’s aggregation or threshold, and an investigation deficiency by changing what the reviewer does, neither of which a documentation project fixes.

The review control documentation template

Two documents per key review control: the design record, completed once and updated when the control changes, and the review log, completed every time the control operates. The design record is what the external auditor asks for at planning; the log is what they sample.

Review control design record

1. Control. Reference; description (full form, naming report, expectation, threshold, investigation, resolution, evidence, deadline, roles); process; accounts and assertions addressed; frequency; classification (IT-dependent manual).

2. Information reviewed. Report or analysis name; source system or workbook; parameters; level of aggregation; IPE test reference; preparer role.

3. Expectation. Source (budget, forecast, prior period adjusted, independent data, model); how it is obtained; independence from the reviewed figures.

4. Precision analysis. Each of the six factors with the evidence; threshold in amount and percentage; tolerable misstatement for the assertion; calculation showing the review would detect a misstatement of that size at the aggregation level used; conclusion (key control for the listed assertions / monitoring control).

5. Investigation and resolution procedure. Steps the reviewer performs on items over threshold; corroboration sources; correction and escalation rules; treatment of uncorrected items.

6. Evidence. Form of evidence (annotated analysis, log, workflow, minutes); retention location; what a tester will find for each instance.

7. Roles. Preparer, reviewer, second reviewer; competence and authority basis; delegation rules.

8. Change history. Changes to threshold, aggregation, information or roles with dates and the re-testing performed.

Review log (per instance)

Period; report instance and parameters; reconciliation of the analysis to the ledger (tie-out reference); list of every item over threshold with amount and percentage; for each item: explanation, corroboration reference, conclusion, correction reference or uncorrected amount; items below threshold investigated at the reviewer’s discretion, if any; reviewer sign-off with date and time relative to close; second-reviewer sign-off; open items carried to the next period.

Worked example: Lakeshore Bancorp’s interest income review

Lakeshore Bancorp, the illustrative $9 billion regional bank used across this site, carries a monthly interest income variance review as a key control over the completeness, accuracy and cut-off of interest income, which at roughly $315 million a year is the largest line in its income statement. As inherited from the program’s first year, the control read: “The Controller reviews the monthly interest income analysis and investigates significant variances.” The external auditor had accepted it for two years and, in the third, asked the precision questions, and the program office could not answer them. The review was performed on a total-bank comparison to budget with an unwritten threshold the Controller described as “anything that looks off, usually a couple of percent”. At 3 percent of monthly interest income, the effective threshold was about $790,000 on a single month and, on the annual figure, roughly $9.5 million; overall materiality was $6.5 million. The review could not, as designed, detect a material misstatement in the assertion it was listed against, and the program recorded a design deficiency at interim, in the same year the deferred-fee spreadsheet error described in the deficiency evaluation guide was found by the auditor rather than by any review.

The redesign, implemented for the third quarter, followed the anatomy. The information became a portfolio-level analysis (nine loan portfolios and four securities and cash categories) generated from the general ledger, with the analysis reconciled to the ledger’s interest income total each month and tested as IPE. The expectation became a computed one: prior month adjusted for the change in average balance and yield taken from the loan system, alongside budget, so that a variance could be decomposed into volume, rate and residual, and the residual was the thing reviewed. The threshold became $250,000 or 2 percent at portfolio level, which the precision calculation showed would surface any single-portfolio misstatement above about $500,000 and every misstatement above $4 million in any combination of portfolios. The investigation procedure required corroboration to the loan system’s yield and balance data or to a specific transaction, recorded in a review log with the explanation, the corroboration reference and the conclusion per item. The deadline became eight business days after month-end, before the close was finalized, and the Chief Accounting Officer reviewed the log and the uncorrected-items list quarterly.

ElementBeforeAfterEffect on precision
InformationTotal-bank interest income vs budgetThirteen portfolios and categories vs budget and vs prior month adjusted for balance and yieldAggregation level at which a $500,000 error is visible
ExpectationBudget onlyComputed expectation from independent balance and yield data plus budgetResidual variance isolates errors from business change
ThresholdUnwritten, roughly 3 percent of the total$250,000 or 2 percent per portfolioEffective detection threshold from about $9.5 million annual to about $500,000
InvestigationPreparer’s explanation acceptedCorroborated to loan system data or a transaction; logged per itemExplanation cannot conceal an error
EvidenceSignature on the analysisReview log with items, explanations, corroboration, conclusions, correctionsTestable without inquiry
Timing and oversightUndated; no second reviewEight business days; quarterly CAO review of the logErrors corrected before finalization; override risk addressed

Internal audit tested the redesigned control for the fourth quarter as part of the year-end program: three monthly instances, every element of the test program, with re-performance of the corroboration on eleven of the 23 flagged items across the three months. The re-performance agreed on ten; on the eleventh, a $310,000 variance in the commercial real estate portfolio explained as “rate reset timing”, the tester’s recomputation from the loan system showed the reset effect was $190,000 and the remainder was a $120,000 accrual error, which the Controller had accepted on the preparer’s explanation. The exception was recorded, the error corrected, and the control concluded as operating effectively with one exception in a sample of 23 items that was evaluated as an isolated corroboration lapse, because the other ten re-performances agreed and the lapse was in the first month of the new procedure. The deficiency at interim was concluded as remediated as of year-end, with the redesigned control having operated for three months, which met the program’s own working rule for a monthly control. The external auditor re-performed the precision calculation, tested two months, and concurred.

Common mistakes

Describing the review without the report, the expectation, the threshold or the deadline. Setting the threshold as a percentage of a large base so that material amounts fall below it. Reviewing at total-company aggregation and asserting precision from the reviewer’s seniority. Letting the preparer write the explanations before the reviewer looks. Accepting explanations without corroboration and calling the file evidence. Treating a signature as evidence of review, or emails as a review log. Testing the reviewer and not the report, or the report and not the reviewer. Concluding on the review by inquiry because the file is thin. Remediating a precision deficiency with a documentation project. And listing the same review as the compensating control for six unrelated deficiencies, which the SOX scoping guide and the deficiency evaluation guide both warn against, because a review’s precision is spent once. The design versus operating effectiveness guide covers the two-stage test the walkthrough and the sample implement here, and the audit evidence guide the evidence hierarchy that puts re-performance above inspection and inspection above inquiry.

Comments

Leave a Reply

Discover more from internalauditguide.com

Subscribe now to keep reading and get access to the full archive.

Continue reading