Internal controls are the most discussed and least defined things in a company. Executives talk about “having controls”, auditors talk about “testing controls”, regulators demand “effective controls”, and in most meetings the three groups mean different things. This guide gives the word a precise meaning, sorts controls into the types that matter for testing, explains who tests them and why several different people test the same control for different reasons, sets out the methods and sample sizes that make a test worth relying on, and follows a real set of controls through a worked example so that the abstractions land on something concrete. It is written for people new to the subject and for practitioners who need a reference they can hand to a process owner without embarrassment.
This guide was rewritten in September 2026 around the COSO Internal Control framework as it stands, the IIA’s Three Lines Model in its 2020 form, and the Global Internal Audit Standards that govern the third line’s testing since January 2025. The worked example is MidState Beverage’s route cash process, the three-state distributor that appears across this site, because cash handled by three hundred drivers is the kind of process where every type of control, every line, and every testing method shows up in a single year. For the framework in full, the COSO 17 principles guide and the COSO framework guide go deeper; for the relationship between the second and third lines, the internal audit vs. compliance guide.
In this guide
- What an internal control is, precisely
- The types of control, and why the type decides the test
- Who tests controls: the Three Lines Model and the people outside it
- Testing methods: design, operation, and the four procedures
- Sample sizes and what they buy
- Worked example: MidState’s route cash controls, line by line
- When a control fails: deficiencies, ratings, and what happens next
- Common mistakes
What an internal control is, precisely
COSO’s 2013 framework, which remains the reference in nearly every jurisdiction, defines internal control as a process, effected by an entity’s board of directors, management, and other personnel, designed to provide reasonable assurance regarding the achievement of objectives relating to operations, reporting, and compliance. Four words in that definition do most of the work. It is a process, not a document: a policy nobody follows is not a control. It is effected by people, which means every control has an operator and the operator’s competence and incentives are part of the control. It provides reasonable assurance, not certainty, so a control that fails occasionally is not necessarily a failed control. And it serves objectives, so a control is only meaningful relative to the risk it addresses; “we reconcile the account monthly” is a control only if you can say which misstatement or loss the reconciliation is there to catch.
An individual control, then, is a specific action, performed by a specific person or system, at a specific frequency, that reduces a specific risk to an acceptable level, and that leaves evidence it was performed. Written that way, most “controls” in a typical process narrative turn out to be fragments: an action without an owner, a review without a stated purpose, a system setting nobody can name. The discipline of writing controls in the full form is what the what is a control guide is about, and the risk and control matrix that results is the document every tester starts from; the free RCM Workbench on this site builds one.
COSO organizes the whole system into five components, and the components matter for testing because they tell you where a failure actually lives. A control activity that keeps failing may be a control environment problem, if the operator’s manager does not care, or an information problem, if the report the operator reviews is wrong. The table sets out the five with their seventeen principles summarized, what each looks like in a company of MidState’s size, and who owns it.
| COSO component | What it covers (principles) | What it looks like at a mid-sized distributor | Who owns it | How it is tested |
|---|---|---|---|---|
| Control environment | Integrity and ethics, board oversight, structure and authority, competence, accountability (principles 1 to 5) | A code of conduct people have read, a board that meets and asks questions, a depot manager who is held to account for a reconciliation | The board and the executive team | Entity-level assessment: interviews, document review, and the pattern of exceptions elsewhere; the culture audit guide covers the method |
| Risk assessment | Objectives specified, risks identified and analyzed, fraud risk considered, change identified (principles 6 to 9) | An annual risk register that the executive team actually updates; a fraud risk assessment that names route cash | Management, with the board’s review | Review of the register against what actually went wrong; the risk assessment guide |
| Control activities | Selected and developed to mitigate risk, technology controls, policies and procedures deployed (principles 10 to 12) | Settlement reconciliations, deposit verification, variance approval, system access rules | Process owners | Tests of design and operation over samples; the bulk of this guide |
| Information and communication | Relevant quality information, internal communication, external communication (principles 13 to 15) | Reports that are complete and accurate, a hotline, customer statements that go out | Management, with IT and finance | Tests of report completeness and accuracy, the IPE testing guide |
| Monitoring activities | Ongoing and separate evaluations, deficiencies evaluated and communicated (principles 16 and 17) | Depot dashboards, the controller’s monthly review, internal audit’s own work, an issue log that closes on evidence | Management for ongoing monitoring; internal audit for separate evaluations | Review of whether monitoring found what audit found first |
The types of control, and why the type decides the test
Controls are classified several ways at once, and each classification answers a different testing question. Timing, whether the control prevents, detects, or corrects, decides what evidence of failure would look like. Nature, whether it is manual, automated, or a manual review of system output, decides how many items you test and whether you need to test the system underneath. Level, whether it operates across the entity or inside one process, decides whether a sample makes sense at all. Importance, whether it is a key control that the objective depends on or one of several that overlap, decides whether it gets tested every year. The table puts the classifications side by side with MidState examples, and the linked guides expand each.
| Classification | Types | MidState example | What the type changes about the test |
|---|---|---|---|
| Timing relative to the error | Preventive, detective, corrective; see the preventive, detective, corrective guide | Preventive: the handheld will not close a route without a cash count. Detective: the depot’s daily settlement reconciliation. Corrective: the variance investigation and re-deposit process | Preventive controls are tested by attempting the prohibited action or inspecting the configuration; detective controls are tested by checking that exceptions were found and acted on; corrective controls by tracing an exception to its resolution |
| Nature | Manual; automated; IT-dependent manual (a person reviewing a system report); see the management review controls guide | Manual: the depot manager signs the deposit listing. Automated: the ERP posts a variance flag above $50. IT-dependent manual: the controller reviews the monthly variance report | Automated controls can be tested with one item plus the general controls over the system; manual controls need a sample sized to their frequency; IT-dependent manual controls need both the review and the report’s completeness tested |
| Level | Entity-level; process- or transaction-level | Entity: the code of conduct and the audit committee’s oversight. Transaction: every one of 37,400 settlements reconciled | Entity-level controls are assessed rather than sampled; transaction-level controls are sampled from a defined population |
| Importance | Key controls that the objective depends on; secondary or compensating controls that overlap | Key: independent settlement reconciliation. Compensating: monthly customer statements that would surface a diverted payment | Key controls are tested every cycle; compensating controls are tested when a key control fails, to see whether the failure was caught downstream |
| Technology layer | IT general controls over access, change, and operations; application controls inside the software; see the ITGC vs. application controls guide | ITGC: who can change the $50 variance threshold. Application: the threshold itself and the flag it raises | Application controls are only reliable if the general controls over the system are; test the general controls first or the application test proves nothing |
| Objective served | Operations, reporting, compliance | Operations: routes settled the same day. Reporting: cash in the ledger equals cash in the bank. Compliance: excise volumes reconciled to returns | Decides which tester cares: the external auditor tests reporting controls; internal audit tests all three; compliance monitors the third |
| Separation of duties | A special case: a design property rather than a single control; see the segregation of duties guide | The driver who collects cash does not reconcile it; the person who overrides a variance is not the person whose variance it is | Tested by examining who holds which access and who actually performed which step, across the whole population rather than a sample |
The most consequential classification in daily work is nature, because it decides the sample. An automated control that operated correctly once, in a system whose change and access controls are sound, operated correctly every time, so the test is one item and the general controls. A manual control performed by a person can be performed well on Monday and badly on Friday, so the test is a sample sized to how often it runs. An IT-dependent manual control is the trap: the depot manager who reviews the variance report every day is only a control if the report is complete and accurate, and testing the manager’s signature without testing the report is the single most common error in control testing, common enough that the IPE guide exists to prevent it.
Who tests controls: the Three Lines Model and the people outside it
The same control can be tested by five different parties in one year, and the reason is not waste. Each tester answers a different question for a different audience. The IIA’s Three Lines Model, revised in 2020 to describe roles rather than organizational boxes, places the first three: management operating and checking its own controls, second-line functions such as risk and compliance monitoring and challenging them, and internal audit providing independent assurance to the board. Outside the model sit the external auditor, whose opinion on internal control over financial reporting under Section 404 of the Sarbanes-Oxley Act goes to shareholders, the regulators who examine controls in regulated industries, and the service auditors who report on controls at outsourced providers. The table sets out who tests, for whom, under which standard, and how far the others can rely on the result.
| Tester | Question answered | Audience | Standard applied | Independence from the control | Who relies on it, and how far |
|---|---|---|---|---|---|
| First line: process owners and their managers | Did my control operate today, this week, this month? | The owner and their line management | The company’s own procedures; management self-assessment programs; SOX 302 sub-certifications | None; the operator or their manager | Second and third lines use it for scoping; nobody treats it as assurance |
| Second line: risk, compliance, internal control, quality functions | Are the first line’s controls designed and operating as the policy requires? | Executive management and management committees | The function’s own monitoring program; regulatory expectations for compliance testing; the SOX program office’s testing for 404(a) | Separate from the operator but part of management; often designed the control being tested | Internal audit may rely partially after assessing objectivity, competence, and method; the external auditor may use management’s 404 testing under its own standards |
| Third line: internal audit | Can the board rely on management’s account of its controls? | The audit committee, then management | The Global Internal Audit Standards; the function’s methodology | Independent of management by charter; reports functionally to the board | The external auditor may use internal audit’s work under PCAOB AS 2605 and the AICPA equivalent, after evaluating competence and objectivity; regulators rely on it heavily and examine it |
| External auditor | Are the controls over financial reporting effective, and can the financial statements be relied on? | Shareholders, through the audit report; the audit committee | PCAOB AS 2201 for integrated audits under SOX 404(b); AICPA and international standards otherwise | Independent of the company under professional and SEC rules | The market; the audit committee; not internal audit, which draws its own conclusions |
| Regulators and examiners | Does the institution meet the regulatory expectation for control? | The regulator, then the board | Examination manuals and supervisory guidance | Independent of the company | The board; management, which must remediate findings; internal audit, which validates remediation |
| Service auditors | Are the controls at an outsourced provider designed and operating? | The provider’s customers and their auditors | SOC 1 under SSAE 18 for financial reporting controls; SOC 2 for security and related criteria | Independent of the provider | User entities and their internal and external auditors, subject to the complementary user controls the report lists; the third-party guide covers the reliance rules |
Two rules govern the traffic between the testers. The first is that reliance flows toward independence, not away from it: internal audit can use the second line’s testing after assessing it, the external auditor can use internal audit’s work after assessing it, but nobody can use first-line self-assessment as assurance, and no amount of downstream testing makes an operator’s checklist independent. The second is that reliance never transfers responsibility: the external auditor who uses internal audit’s work still owns the opinion, and the CAE who relies on compliance’s monitoring still owns the conclusion reported to the committee. GIAS Standard 9.5 states the third line’s version of that rule, and the Domain III guide covers the independence on which the whole chain rests.
Testing methods: design, operation, and the four procedures
Every control test answers two questions in order. Is the control designed so that, if it operates as described, it would actually address the risk? And did it operate as described throughout the period? A test of design is performed once per control, usually through a walkthrough of a single transaction from start to finish with the people who perform each step, and it fails when the control could operate perfectly and still miss the error: a reconciliation performed by the person who handles the cash, a review threshold set above the size of loss that matters, a report that omits the transactions most likely to be wrong. A test of operating effectiveness is performed over a sample of instances across the period and fails when the designed control was not performed, was performed late, was performed by the wrong person, or was performed without the evidence that would let anyone know. The design vs. operating effectiveness guide works through both with examples, and the walkthrough template shows the design test’s documentation.
Within those two questions there are four procedures, and they are listed in ascending order of strength for a reason: each one produces better evidence than the last and costs more to perform. The choice among them is the tester’s main judgment, and the standard against which it is judged is whether the evidence would persuade someone who was not there. The audit evidence guide sets out that standard.
| Procedure | What it involves | What it proves | Sufficient on its own when | Weakness |
|---|---|---|---|---|
| Inquiry | Asking the operator or their manager how the control works and whether it was performed | What people believe happens | Never, for operating effectiveness; useful for design understanding and for corroborating other evidence | People describe the procedure, not the practice; the answer is the same whether the control ran or not |
| Observation | Watching the control being performed | That it operated on the day you watched | Controls that leave no evidence, such as a physical count or a supervisor’s presence, tested at several unannounced points | Proves one instance; people perform well while watched |
| Inspection | Examining the documents, records, or system evidence that the control was performed: the signed listing, the reviewer’s comments, the system log, the ticket | That the evidence of performance exists for the items examined | Most manual controls with a documented trail, sampled across the period; automated controls where configuration can be inspected | A signature proves a signature; it does not prove the reviewer looked at anything, which is why review controls need evidence of what the reviewer did |
| Re-performance | Doing the control again independently and comparing the result: redoing the reconciliation, recomputing the variance, re-running the access query | That the control, performed correctly, produces the right answer, and whether the operator’s result matched | Key controls, high-risk controls, and any control whose evidence is thin; the standard for reliance decisions | Expensive; needs the underlying data; can drift into the tester performing the control |
Two procedures that cut across the four deserve their own mention. Data analytics over the full population, such as recalculating every settlement’s variance and listing every override with its approver, is re-performance at scale, and where the data exists it replaces sampling for the parts of the control it can reach; the journal entry analytics guide shows the pattern. Testing information produced by the entity, meaning proving that the report the control operator relied on was complete and accurate, is a prerequisite rather than a procedure, and a review control tested without it is a review of an untested report.
Sample sizes and what they buy
Sample sizes for control testing are conventions rather than statistics in most functions, and the conventions are defensible because they derive from an attribute-sampling model with a low expected error rate and a tolerable rate around ten percent at high confidence. The table gives the sizes this site treats as standard, explained in the sample size guide, with the population each is drawn from and what a single exception does to the conclusion. The free mySampler tool draws the samples.
| Control frequency | Population in a year | Standard sample | Higher-risk sample | What one exception means |
|---|---|---|---|---|
| Annual | 1 | 1 | 1, plus re-performance | The control failed for the year |
| Quarterly | 4 | 2 | 4 | Test all four; a second exception is a failed control |
| Monthly | 12 | 2 to 4 | 6 | Expand to all twelve and evaluate the rate |
| Weekly | 52 | 5 to 10 | 15 | Expand the sample; one exception in five is a failed control, one in fifteen is a judgment |
| Daily | 250 or so | 20 to 25 | 40 | Expand by at least the original size; evaluate the exception’s cause before the rate |
| Many times a day | Thousands | 25 to 40 | 60 | Same; consider full-population analytics instead of expanding |
| Automated | Any | 1, plus the general controls over the system | 1 per configuration variant | The control or the general controls failed; there is no rate to evaluate |
Three points of practice. The sample is only as good as the population it came from, so the first test in every file is the completeness of the list you sampled: a settlement population pulled from the ERP that omits the two acquired depots still on spreadsheets is not a population, and the random sampling guide covers how to build and document one. Exceptions are investigated before they are counted, because an exception that turns out to be a documentation lapse and one that turns out to be a missing control get the same tally and different findings. And a sample designed to conclude on operating effectiveness expects zero exceptions; the convention sizes assume it, and the moment an exception appears the tester is no longer estimating a rate but deciding whether the control can be relied on at all.
Worked example: MidState’s route cash controls, line by line
MidState Beverage collects about $31 million a year in cash and checks from roughly 4,200 smaller customers through three hundred delivery routes and twelve depots, on a 2013 route-accounting module plus depot spreadsheets, with two acquired distributors not yet on the ERP. In FY26 a Dayton driver diverted $18,400 over five months before a customer complaint exposed it. The FY27-01 route cash audit that followed tested every control in the settlement process and rated the engagement Unsatisfactory with five findings. The table lists the controls, their types, which line operates and monitors each, what internal audit did, and what it found. MidState has no second-line function; the controller’s finance review is management-level monitoring rather than an independent second line, and the gap shows in the results.
| Control | Type | First line: who operates it | Management monitoring | Third line: what internal audit tested | Result |
|---|---|---|---|---|---|
| Handheld will not close a route without a cash count and a settlement record | Preventive, automated application control | Drivers, through the device; IT maintains the configuration | None beyond IT’s change process | Configuration inspected; change log for the handheld software reviewed for the year; one route closure re-performed per device version | Operating; but sync failures between handheld and ERP were not monitored, so settlements could sit unposted for days (finding 3) |
| Daily settlement reconciliation: cash and checks counted against the handheld settlement, signed by a clerk and reviewed by the depot manager | Detective, manual | Depot clerk prepares; depot manager reviews | Controller’s monthly variance report | 60 reconciliations stratified across all twelve depots, inspected and re-performed; preparer and reviewer identity checked | 14 of 60 failed; at nine depots the reviewer was the same person who prepared, or had handled the cash (finding 1) |
| Variance override approval: any settlement variance above $50 requires a manager’s approval in the ERP | Detective, IT-dependent manual | Depot managers approve; the ERP records approver | None; the override log was not reviewed above depot level | Full-population analytic of 37,400 settlements: every override matched to its approver and the approver’s role | 1,412 overrides were approved by the user who created the variance (finding 2); segregation of duties failed by design, not by exception |
| Deposit listing prepared from the ERP and agreed to the bank deposit | Detective, IT-dependent manual | Depot clerk prepares; depot manager signs | Controller’s bank reconciliation | 70 deposits selected by monetary unit sampling, agreed from listing to bank and from bank to listing; the source of each listing inspected | Deposits agreed to bank; but at seven depots the listing was re-keyed into a spreadsheet outside the ERP, breaking the trail from settlement to deposit (finding 4) |
| Monthly customer statements mailed or emailed to every cash-paying customer | Detective, compensating: the control that would let a customer notice a diverted payment | Accounts receivable, from the ERP customer master | None | Customer master compared to the statement run; 20 customers selected for targeted confirmation from the unstated population | 1,130 of 4,200 customers were flagged to receive no statement, mostly acquired-distributor accounts (finding 5); confirmed customers had nothing to compare a payment against, which is how the Dayton loss ran for five months |
| Controller’s monthly variance review across depots | Monitoring, IT-dependent manual | Controller | CFO reads the summary | Report completeness tested against the settlement population; reviewer’s evidence of follow-up inspected for six months | Report excluded the two acquired depots, which were not on the ERP; the review was performed but over an incomplete population |
| Segregation of duties across the process: collector, reconciler, approver, depositor | Design property | Everyone in the process | None formal | ERP role assignments and actual performers analyzed for every settlement | Failed at nine depots for reconciliation and across the override process; the design assumed staffing the depots did not have |
What the example shows about the lines is that a missing second line is a hole in the middle of the picture. Each depot’s first-line controls existed on paper and were performed after a fashion, the controller’s monitoring existed but ran on incomplete data, and nobody between the depots and internal audit was looking across all twelve at once, so the same weakness repeated at nine of them without anyone noticing until an auditor sampled all twelve. The eleven management actions included a compliance coordinator role in operations, described in the internal audit vs. compliance guide, and a monthly override and reconciliation dashboard that gives management the cross-depot view audit had to build itself. The report examples guide shows how findings of this kind are written.
When a control fails: deficiencies, ratings, and what happens next
A failed test produces a deficiency, and what happens next depends on which objective the control served. For controls over financial reporting the vocabulary is fixed by the auditing standards: a deficiency exists when a control’s design or operation does not allow management or employees to prevent or detect misstatements on a timely basis; a significant deficiency is less severe than a material weakness but important enough to merit the attention of those responsible for oversight of financial reporting; and a material weakness is a deficiency, or combination of deficiencies, such that there is a reasonable possibility that a material misstatement will not be prevented or detected on a timely basis. Material weaknesses are disclosed publicly in the 404 report, which is why the classification is argued over so hard, and the deficiency evaluation guide walks through the judgment. For operational and compliance controls there is no statutory scale, and internal audit uses its own severity ratings, which the severity ratings guide sets out.
| Step after a failed test | Financial reporting controls (SOX 404) | Operational and compliance controls (internal audit) |
|---|---|---|
| Establish the cause | Design or operation; isolated or systemic; the tester documents both before classifying | Same; root cause is the field that decides whether the fix will hold |
| Consider compensating controls | A compensating control that operates at the right precision can reduce severity; it must itself be tested | Same; MidState’s customer statements were the compensating control, and they had failed too |
| Assess magnitude and likelihood | Potential misstatement against materiality; probability judged as reasonably possible or remote | Potential loss, regulatory, operational, or reputational impact against the rating scale; likelihood judged from the exception rate and the environment |
| Aggregate | Deficiencies in the same account or process are combined; several small ones can be a material weakness together | Findings in the same process are read together for the engagement rating; five related findings made FY27-01 Unsatisfactory |
| Classify | Deficiency, significant deficiency, material weakness; concluded by management and evaluated by the external auditor | Low to critical finding ratings and an overall engagement rating; concluded by the CAE |
| Report | Significant deficiencies and material weaknesses to the audit committee and external auditor; material weaknesses disclosed publicly | All findings to management; high and critical findings and Unsatisfactory engagements to the audit committee |
| Remediate and retest | Management remediates; the control must operate for a sufficient period before year-end to be tested and relied on | Management action plans with owners and dates; internal audit validates on evidence before closure, as the issue validation guide describes |
Common mistakes
| Mistake | What it looks like | Why it matters | Fix |
|---|---|---|---|
| Controls without a risk | A control matrix listing activities with no stated risk each one addresses | Nobody can say whether the control is designed adequately, because adequate means adequate for something | Write every control against the risk it mitigates; delete the ones that mitigate nothing |
| Testing the signature | Inspecting that a review was signed, without evidence of what the reviewer examined or the report’s completeness | A review control with no evidence of review is not operating, however many signatures exist | Require evidence of the review’s content; test the report under it |
| Sampling from an incomplete population | Pulling the sample from the ERP when part of the business runs on spreadsheets | The conclusion covers only the part of the population that was in the system | Prove completeness first; MidState’s acquired depots were the gap |
| Treating self-assessment as assurance | The committee pack reports “controls effective” based on managers’ attestations | The first line cannot be independent of itself | Label self-assessment as such; assurance comes from the third line or an external party |
| Testing automated controls like manual ones | Sampling 25 instances of a system check while never testing who can change its configuration | The 25 items prove nothing the first one did not, and the real risk, an unauthorized change, was untested | One item plus the general controls |
| Counting exceptions before understanding them | “Three of twenty-five failed” reported as a rate without causes | A documentation lapse and a missing control get the same score | Investigate each exception; report cause, then rate |
| Ignoring segregation of duties as a design question | Testing that reconciliations were performed without checking who performed them | A reconciliation by the cash handler is a reconciliation of nothing | Test the performer’s identity and role in every sample item |
| Closing on management’s word | Issues marked remediated when the action plan says so | Findings recur, and the committee learns it from the next audit or the next loss | Validate on evidence; keep the issue open until the control has operated |
| Everyone tests, nobody coordinates | The second line, internal audit, and the external auditor sample the same control in the same quarter with different methods | Auditee fatigue and three different conclusions | A combined assurance map with reliance decisions written down |
The whole subject reduces to a few sentences. A control is a specific action by a specific person or system that reduces a specific risk and leaves evidence. Its type decides how it is tested and how many items the test needs. Several parties test the same control because each answers a different question for a different audience, and reliance flows toward the independent tester without ever transferring responsibility. A failed test is the beginning of a judgment about cause, severity, and remediation, not the end of one. Get those four things right and the rest of control testing is diligence; get any of them wrong and a clean test result means nothing at all.
Related guides
- What is a control
- The COSO 17 principles
- Preventive, detective, and corrective controls
- Test of design vs. operating effectiveness
- Audit sample sizes: 25, 40, 60
- Audit evidence
- Testing information produced by the entity
- Management review controls
- Segregation of duties
- ITGC vs. application controls
- SOX 404 explained
- Control deficiency evaluation
- Internal audit vs. compliance
- Substantive testing for beginners
- Start here
Leave a Reply