,

Internal Controls: What They Are, Who Tests Them, and How

Internal controls are the most discussed and least defined things in a company. Executives talk about “having controls”, auditors talk about “testing controls”, regulators demand “effective controls”, and in most meetings the three groups mean different things. This guide gives the word a precise meaning, sorts controls into the types that matter for testing, explains who tests them and why several different people test the same control for different reasons, sets out the methods and sample sizes that make a test worth relying on, and follows a real set of controls through a worked example so that the abstractions land on something concrete. It is written for people new to the subject and for practitioners who need a reference they can hand to a process owner without embarrassment.

This guide was rewritten in September 2026 around the COSO Internal Control framework as it stands, the IIA’s Three Lines Model in its 2020 form, and the Global Internal Audit Standards that govern the third line’s testing since January 2025. The worked example is MidState Beverage’s route cash process, the three-state distributor that appears across this site, because cash handled by three hundred drivers is the kind of process where every type of control, every line, and every testing method shows up in a single year. For the framework in full, the COSO 17 principles guide and the COSO framework guide go deeper; for the relationship between the second and third lines, the internal audit vs. compliance guide.

In this guide

What an internal control is, precisely

COSO’s 2013 framework, which remains the reference in nearly every jurisdiction, defines internal control as a process, effected by an entity’s board of directors, management, and other personnel, designed to provide reasonable assurance regarding the achievement of objectives relating to operations, reporting, and compliance. Four words in that definition do most of the work. It is a process, not a document: a policy nobody follows is not a control. It is effected by people, which means every control has an operator and the operator’s competence and incentives are part of the control. It provides reasonable assurance, not certainty, so a control that fails occasionally is not necessarily a failed control. And it serves objectives, so a control is only meaningful relative to the risk it addresses; “we reconcile the account monthly” is a control only if you can say which misstatement or loss the reconciliation is there to catch.

An individual control, then, is a specific action, performed by a specific person or system, at a specific frequency, that reduces a specific risk to an acceptable level, and that leaves evidence it was performed. Written that way, most “controls” in a typical process narrative turn out to be fragments: an action without an owner, a review without a stated purpose, a system setting nobody can name. The discipline of writing controls in the full form is what the what is a control guide is about, and the risk and control matrix that results is the document every tester starts from; the free RCM Workbench on this site builds one.

COSO organizes the whole system into five components, and the components matter for testing because they tell you where a failure actually lives. A control activity that keeps failing may be a control environment problem, if the operator’s manager does not care, or an information problem, if the report the operator reviews is wrong. The table sets out the five with their seventeen principles summarized, what each looks like in a company of MidState’s size, and who owns it.

COSO componentWhat it covers (principles)What it looks like at a mid-sized distributorWho owns itHow it is tested
Control environmentIntegrity and ethics, board oversight, structure and authority, competence, accountability (principles 1 to 5)A code of conduct people have read, a board that meets and asks questions, a depot manager who is held to account for a reconciliationThe board and the executive teamEntity-level assessment: interviews, document review, and the pattern of exceptions elsewhere; the culture audit guide covers the method
Risk assessmentObjectives specified, risks identified and analyzed, fraud risk considered, change identified (principles 6 to 9)An annual risk register that the executive team actually updates; a fraud risk assessment that names route cashManagement, with the board’s reviewReview of the register against what actually went wrong; the risk assessment guide
Control activitiesSelected and developed to mitigate risk, technology controls, policies and procedures deployed (principles 10 to 12)Settlement reconciliations, deposit verification, variance approval, system access rulesProcess ownersTests of design and operation over samples; the bulk of this guide
Information and communicationRelevant quality information, internal communication, external communication (principles 13 to 15)Reports that are complete and accurate, a hotline, customer statements that go outManagement, with IT and financeTests of report completeness and accuracy, the IPE testing guide
Monitoring activitiesOngoing and separate evaluations, deficiencies evaluated and communicated (principles 16 and 17)Depot dashboards, the controller’s monthly review, internal audit’s own work, an issue log that closes on evidenceManagement for ongoing monitoring; internal audit for separate evaluationsReview of whether monitoring found what audit found first

The types of control, and why the type decides the test

Controls are classified several ways at once, and each classification answers a different testing question. Timing, whether the control prevents, detects, or corrects, decides what evidence of failure would look like. Nature, whether it is manual, automated, or a manual review of system output, decides how many items you test and whether you need to test the system underneath. Level, whether it operates across the entity or inside one process, decides whether a sample makes sense at all. Importance, whether it is a key control that the objective depends on or one of several that overlap, decides whether it gets tested every year. The table puts the classifications side by side with MidState examples, and the linked guides expand each.

ClassificationTypesMidState exampleWhat the type changes about the test
Timing relative to the errorPreventive, detective, corrective; see the preventive, detective, corrective guidePreventive: the handheld will not close a route without a cash count. Detective: the depot’s daily settlement reconciliation. Corrective: the variance investigation and re-deposit processPreventive controls are tested by attempting the prohibited action or inspecting the configuration; detective controls are tested by checking that exceptions were found and acted on; corrective controls by tracing an exception to its resolution
NatureManual; automated; IT-dependent manual (a person reviewing a system report); see the management review controls guideManual: the depot manager signs the deposit listing. Automated: the ERP posts a variance flag above $50. IT-dependent manual: the controller reviews the monthly variance reportAutomated controls can be tested with one item plus the general controls over the system; manual controls need a sample sized to their frequency; IT-dependent manual controls need both the review and the report’s completeness tested
LevelEntity-level; process- or transaction-levelEntity: the code of conduct and the audit committee’s oversight. Transaction: every one of 37,400 settlements reconciledEntity-level controls are assessed rather than sampled; transaction-level controls are sampled from a defined population
ImportanceKey controls that the objective depends on; secondary or compensating controls that overlapKey: independent settlement reconciliation. Compensating: monthly customer statements that would surface a diverted paymentKey controls are tested every cycle; compensating controls are tested when a key control fails, to see whether the failure was caught downstream
Technology layerIT general controls over access, change, and operations; application controls inside the software; see the ITGC vs. application controls guideITGC: who can change the $50 variance threshold. Application: the threshold itself and the flag it raisesApplication controls are only reliable if the general controls over the system are; test the general controls first or the application test proves nothing
Objective servedOperations, reporting, complianceOperations: routes settled the same day. Reporting: cash in the ledger equals cash in the bank. Compliance: excise volumes reconciled to returnsDecides which tester cares: the external auditor tests reporting controls; internal audit tests all three; compliance monitors the third
Separation of dutiesA special case: a design property rather than a single control; see the segregation of duties guideThe driver who collects cash does not reconcile it; the person who overrides a variance is not the person whose variance it isTested by examining who holds which access and who actually performed which step, across the whole population rather than a sample

The most consequential classification in daily work is nature, because it decides the sample. An automated control that operated correctly once, in a system whose change and access controls are sound, operated correctly every time, so the test is one item and the general controls. A manual control performed by a person can be performed well on Monday and badly on Friday, so the test is a sample sized to how often it runs. An IT-dependent manual control is the trap: the depot manager who reviews the variance report every day is only a control if the report is complete and accurate, and testing the manager’s signature without testing the report is the single most common error in control testing, common enough that the IPE guide exists to prevent it.

Who tests controls: the Three Lines Model and the people outside it

The same control can be tested by five different parties in one year, and the reason is not waste. Each tester answers a different question for a different audience. The IIA’s Three Lines Model, revised in 2020 to describe roles rather than organizational boxes, places the first three: management operating and checking its own controls, second-line functions such as risk and compliance monitoring and challenging them, and internal audit providing independent assurance to the board. Outside the model sit the external auditor, whose opinion on internal control over financial reporting under Section 404 of the Sarbanes-Oxley Act goes to shareholders, the regulators who examine controls in regulated industries, and the service auditors who report on controls at outsourced providers. The table sets out who tests, for whom, under which standard, and how far the others can rely on the result.

TesterQuestion answeredAudienceStandard appliedIndependence from the controlWho relies on it, and how far
First line: process owners and their managersDid my control operate today, this week, this month?The owner and their line managementThe company’s own procedures; management self-assessment programs; SOX 302 sub-certificationsNone; the operator or their managerSecond and third lines use it for scoping; nobody treats it as assurance
Second line: risk, compliance, internal control, quality functionsAre the first line’s controls designed and operating as the policy requires?Executive management and management committeesThe function’s own monitoring program; regulatory expectations for compliance testing; the SOX program office’s testing for 404(a)Separate from the operator but part of management; often designed the control being testedInternal audit may rely partially after assessing objectivity, competence, and method; the external auditor may use management’s 404 testing under its own standards
Third line: internal auditCan the board rely on management’s account of its controls?The audit committee, then managementThe Global Internal Audit Standards; the function’s methodologyIndependent of management by charter; reports functionally to the boardThe external auditor may use internal audit’s work under PCAOB AS 2605 and the AICPA equivalent, after evaluating competence and objectivity; regulators rely on it heavily and examine it
External auditorAre the controls over financial reporting effective, and can the financial statements be relied on?Shareholders, through the audit report; the audit committeePCAOB AS 2201 for integrated audits under SOX 404(b); AICPA and international standards otherwiseIndependent of the company under professional and SEC rulesThe market; the audit committee; not internal audit, which draws its own conclusions
Regulators and examinersDoes the institution meet the regulatory expectation for control?The regulator, then the boardExamination manuals and supervisory guidanceIndependent of the companyThe board; management, which must remediate findings; internal audit, which validates remediation
Service auditorsAre the controls at an outsourced provider designed and operating?The provider’s customers and their auditorsSOC 1 under SSAE 18 for financial reporting controls; SOC 2 for security and related criteriaIndependent of the providerUser entities and their internal and external auditors, subject to the complementary user controls the report lists; the third-party guide covers the reliance rules

Two rules govern the traffic between the testers. The first is that reliance flows toward independence, not away from it: internal audit can use the second line’s testing after assessing it, the external auditor can use internal audit’s work after assessing it, but nobody can use first-line self-assessment as assurance, and no amount of downstream testing makes an operator’s checklist independent. The second is that reliance never transfers responsibility: the external auditor who uses internal audit’s work still owns the opinion, and the CAE who relies on compliance’s monitoring still owns the conclusion reported to the committee. GIAS Standard 9.5 states the third line’s version of that rule, and the Domain III guide covers the independence on which the whole chain rests.

Testing methods: design, operation, and the four procedures

Every control test answers two questions in order. Is the control designed so that, if it operates as described, it would actually address the risk? And did it operate as described throughout the period? A test of design is performed once per control, usually through a walkthrough of a single transaction from start to finish with the people who perform each step, and it fails when the control could operate perfectly and still miss the error: a reconciliation performed by the person who handles the cash, a review threshold set above the size of loss that matters, a report that omits the transactions most likely to be wrong. A test of operating effectiveness is performed over a sample of instances across the period and fails when the designed control was not performed, was performed late, was performed by the wrong person, or was performed without the evidence that would let anyone know. The design vs. operating effectiveness guide works through both with examples, and the walkthrough template shows the design test’s documentation.

Within those two questions there are four procedures, and they are listed in ascending order of strength for a reason: each one produces better evidence than the last and costs more to perform. The choice among them is the tester’s main judgment, and the standard against which it is judged is whether the evidence would persuade someone who was not there. The audit evidence guide sets out that standard.

ProcedureWhat it involvesWhat it provesSufficient on its own whenWeakness
InquiryAsking the operator or their manager how the control works and whether it was performedWhat people believe happensNever, for operating effectiveness; useful for design understanding and for corroborating other evidencePeople describe the procedure, not the practice; the answer is the same whether the control ran or not
ObservationWatching the control being performedThat it operated on the day you watchedControls that leave no evidence, such as a physical count or a supervisor’s presence, tested at several unannounced pointsProves one instance; people perform well while watched
InspectionExamining the documents, records, or system evidence that the control was performed: the signed listing, the reviewer’s comments, the system log, the ticketThat the evidence of performance exists for the items examinedMost manual controls with a documented trail, sampled across the period; automated controls where configuration can be inspectedA signature proves a signature; it does not prove the reviewer looked at anything, which is why review controls need evidence of what the reviewer did
Re-performanceDoing the control again independently and comparing the result: redoing the reconciliation, recomputing the variance, re-running the access queryThat the control, performed correctly, produces the right answer, and whether the operator’s result matchedKey controls, high-risk controls, and any control whose evidence is thin; the standard for reliance decisionsExpensive; needs the underlying data; can drift into the tester performing the control

Two procedures that cut across the four deserve their own mention. Data analytics over the full population, such as recalculating every settlement’s variance and listing every override with its approver, is re-performance at scale, and where the data exists it replaces sampling for the parts of the control it can reach; the journal entry analytics guide shows the pattern. Testing information produced by the entity, meaning proving that the report the control operator relied on was complete and accurate, is a prerequisite rather than a procedure, and a review control tested without it is a review of an untested report.

Sample sizes and what they buy

Sample sizes for control testing are conventions rather than statistics in most functions, and the conventions are defensible because they derive from an attribute-sampling model with a low expected error rate and a tolerable rate around ten percent at high confidence. The table gives the sizes this site treats as standard, explained in the sample size guide, with the population each is drawn from and what a single exception does to the conclusion. The free mySampler tool draws the samples.

Control frequencyPopulation in a yearStandard sampleHigher-risk sampleWhat one exception means
Annual111, plus re-performanceThe control failed for the year
Quarterly424Test all four; a second exception is a failed control
Monthly122 to 46Expand to all twelve and evaluate the rate
Weekly525 to 1015Expand the sample; one exception in five is a failed control, one in fifteen is a judgment
Daily250 or so20 to 2540Expand by at least the original size; evaluate the exception’s cause before the rate
Many times a dayThousands25 to 4060Same; consider full-population analytics instead of expanding
AutomatedAny1, plus the general controls over the system1 per configuration variantThe control or the general controls failed; there is no rate to evaluate

Three points of practice. The sample is only as good as the population it came from, so the first test in every file is the completeness of the list you sampled: a settlement population pulled from the ERP that omits the two acquired depots still on spreadsheets is not a population, and the random sampling guide covers how to build and document one. Exceptions are investigated before they are counted, because an exception that turns out to be a documentation lapse and one that turns out to be a missing control get the same tally and different findings. And a sample designed to conclude on operating effectiveness expects zero exceptions; the convention sizes assume it, and the moment an exception appears the tester is no longer estimating a rate but deciding whether the control can be relied on at all.

Worked example: MidState’s route cash controls, line by line

MidState Beverage collects about $31 million a year in cash and checks from roughly 4,200 smaller customers through three hundred delivery routes and twelve depots, on a 2013 route-accounting module plus depot spreadsheets, with two acquired distributors not yet on the ERP. In FY26 a Dayton driver diverted $18,400 over five months before a customer complaint exposed it. The FY27-01 route cash audit that followed tested every control in the settlement process and rated the engagement Unsatisfactory with five findings. The table lists the controls, their types, which line operates and monitors each, what internal audit did, and what it found. MidState has no second-line function; the controller’s finance review is management-level monitoring rather than an independent second line, and the gap shows in the results.

ControlTypeFirst line: who operates itManagement monitoringThird line: what internal audit testedResult
Handheld will not close a route without a cash count and a settlement recordPreventive, automated application controlDrivers, through the device; IT maintains the configurationNone beyond IT’s change processConfiguration inspected; change log for the handheld software reviewed for the year; one route closure re-performed per device versionOperating; but sync failures between handheld and ERP were not monitored, so settlements could sit unposted for days (finding 3)
Daily settlement reconciliation: cash and checks counted against the handheld settlement, signed by a clerk and reviewed by the depot managerDetective, manualDepot clerk prepares; depot manager reviewsController’s monthly variance report60 reconciliations stratified across all twelve depots, inspected and re-performed; preparer and reviewer identity checked14 of 60 failed; at nine depots the reviewer was the same person who prepared, or had handled the cash (finding 1)
Variance override approval: any settlement variance above $50 requires a manager’s approval in the ERPDetective, IT-dependent manualDepot managers approve; the ERP records approverNone; the override log was not reviewed above depot levelFull-population analytic of 37,400 settlements: every override matched to its approver and the approver’s role1,412 overrides were approved by the user who created the variance (finding 2); segregation of duties failed by design, not by exception
Deposit listing prepared from the ERP and agreed to the bank depositDetective, IT-dependent manualDepot clerk prepares; depot manager signsController’s bank reconciliation70 deposits selected by monetary unit sampling, agreed from listing to bank and from bank to listing; the source of each listing inspectedDeposits agreed to bank; but at seven depots the listing was re-keyed into a spreadsheet outside the ERP, breaking the trail from settlement to deposit (finding 4)
Monthly customer statements mailed or emailed to every cash-paying customerDetective, compensating: the control that would let a customer notice a diverted paymentAccounts receivable, from the ERP customer masterNoneCustomer master compared to the statement run; 20 customers selected for targeted confirmation from the unstated population1,130 of 4,200 customers were flagged to receive no statement, mostly acquired-distributor accounts (finding 5); confirmed customers had nothing to compare a payment against, which is how the Dayton loss ran for five months
Controller’s monthly variance review across depotsMonitoring, IT-dependent manualControllerCFO reads the summaryReport completeness tested against the settlement population; reviewer’s evidence of follow-up inspected for six monthsReport excluded the two acquired depots, which were not on the ERP; the review was performed but over an incomplete population
Segregation of duties across the process: collector, reconciler, approver, depositorDesign propertyEveryone in the processNone formalERP role assignments and actual performers analyzed for every settlementFailed at nine depots for reconciliation and across the override process; the design assumed staffing the depots did not have

What the example shows about the lines is that a missing second line is a hole in the middle of the picture. Each depot’s first-line controls existed on paper and were performed after a fashion, the controller’s monitoring existed but ran on incomplete data, and nobody between the depots and internal audit was looking across all twelve at once, so the same weakness repeated at nine of them without anyone noticing until an auditor sampled all twelve. The eleven management actions included a compliance coordinator role in operations, described in the internal audit vs. compliance guide, and a monthly override and reconciliation dashboard that gives management the cross-depot view audit had to build itself. The report examples guide shows how findings of this kind are written.

When a control fails: deficiencies, ratings, and what happens next

A failed test produces a deficiency, and what happens next depends on which objective the control served. For controls over financial reporting the vocabulary is fixed by the auditing standards: a deficiency exists when a control’s design or operation does not allow management or employees to prevent or detect misstatements on a timely basis; a significant deficiency is less severe than a material weakness but important enough to merit the attention of those responsible for oversight of financial reporting; and a material weakness is a deficiency, or combination of deficiencies, such that there is a reasonable possibility that a material misstatement will not be prevented or detected on a timely basis. Material weaknesses are disclosed publicly in the 404 report, which is why the classification is argued over so hard, and the deficiency evaluation guide walks through the judgment. For operational and compliance controls there is no statutory scale, and internal audit uses its own severity ratings, which the severity ratings guide sets out.

Step after a failed testFinancial reporting controls (SOX 404)Operational and compliance controls (internal audit)
Establish the causeDesign or operation; isolated or systemic; the tester documents both before classifyingSame; root cause is the field that decides whether the fix will hold
Consider compensating controlsA compensating control that operates at the right precision can reduce severity; it must itself be testedSame; MidState’s customer statements were the compensating control, and they had failed too
Assess magnitude and likelihoodPotential misstatement against materiality; probability judged as reasonably possible or remotePotential loss, regulatory, operational, or reputational impact against the rating scale; likelihood judged from the exception rate and the environment
AggregateDeficiencies in the same account or process are combined; several small ones can be a material weakness togetherFindings in the same process are read together for the engagement rating; five related findings made FY27-01 Unsatisfactory
ClassifyDeficiency, significant deficiency, material weakness; concluded by management and evaluated by the external auditorLow to critical finding ratings and an overall engagement rating; concluded by the CAE
ReportSignificant deficiencies and material weaknesses to the audit committee and external auditor; material weaknesses disclosed publiclyAll findings to management; high and critical findings and Unsatisfactory engagements to the audit committee
Remediate and retestManagement remediates; the control must operate for a sufficient period before year-end to be tested and relied onManagement action plans with owners and dates; internal audit validates on evidence before closure, as the issue validation guide describes

Common mistakes

MistakeWhat it looks likeWhy it mattersFix
Controls without a riskA control matrix listing activities with no stated risk each one addressesNobody can say whether the control is designed adequately, because adequate means adequate for somethingWrite every control against the risk it mitigates; delete the ones that mitigate nothing
Testing the signatureInspecting that a review was signed, without evidence of what the reviewer examined or the report’s completenessA review control with no evidence of review is not operating, however many signatures existRequire evidence of the review’s content; test the report under it
Sampling from an incomplete populationPulling the sample from the ERP when part of the business runs on spreadsheetsThe conclusion covers only the part of the population that was in the systemProve completeness first; MidState’s acquired depots were the gap
Treating self-assessment as assuranceThe committee pack reports “controls effective” based on managers’ attestationsThe first line cannot be independent of itselfLabel self-assessment as such; assurance comes from the third line or an external party
Testing automated controls like manual onesSampling 25 instances of a system check while never testing who can change its configurationThe 25 items prove nothing the first one did not, and the real risk, an unauthorized change, was untestedOne item plus the general controls
Counting exceptions before understanding them“Three of twenty-five failed” reported as a rate without causesA documentation lapse and a missing control get the same scoreInvestigate each exception; report cause, then rate
Ignoring segregation of duties as a design questionTesting that reconciliations were performed without checking who performed themA reconciliation by the cash handler is a reconciliation of nothingTest the performer’s identity and role in every sample item
Closing on management’s wordIssues marked remediated when the action plan says soFindings recur, and the committee learns it from the next audit or the next lossValidate on evidence; keep the issue open until the control has operated
Everyone tests, nobody coordinatesThe second line, internal audit, and the external auditor sample the same control in the same quarter with different methodsAuditee fatigue and three different conclusionsA combined assurance map with reliance decisions written down

The whole subject reduces to a few sentences. A control is a specific action by a specific person or system that reduces a specific risk and leaves evidence. Its type decides how it is tested and how many items the test needs. Several parties test the same control because each answers a different question for a different audience, and reliance flows toward the independent tester without ever transferring responsibility. A failed test is the beginning of a judgment about cause, severity, and remediation, not the end of one. Get those four things right and the rest of control testing is diligence; get any of them wrong and a clean test result means nothing at all.

Related guides

Comments

One response to “Internal Controls: What They Are, Who Tests Them, and How”

  1. […] of duties—remain robust. If the government moves toward a formal “directors’ statement on internal controls,” internal audit must systematically test these controls (Section 404-like under US SOX […]

Leave a Reply

Discover more from internalauditguide.com

Subscribe now to keep reading and get access to the full archive.

Continue reading