,

Operational Risk Guide for Internal Auditors (2026)

Operational risk is the risk that never appears on a trading blotter and accounts for most of the losses that end careers: the wire sent to the wrong account, the model that was never revalidated, the vendor whose breach became yours, the control that a reorganization quietly removed. Basel’s definition, the risk of loss resulting from inadequate or failed internal processes, people and systems or from external events, is deliberately broad, and the breadth is why so many operational risk frameworks are wide and shallow: a taxonomy nobody uses, an RCSA that produces the same green ratings every year, a KRI dashboard nobody acts on, and a loss database that records the losses the business chose to report. The framework is only as good as the discipline behind each component, and internal audit is the function whose job is to say whether that discipline exists.

This guide covers operational risk from both sides of the third line: what the risk is, how it is classified and measured, what an operational risk management framework consists of, and how internal audit audits the framework and its use. It gives you the seven Basel event types with working examples, a framework components table with the artifact and the audit test for each, a set of twelve key risk indicators with thresholds, a worked loss-data example, a fifteen-row audit work program, and the failures that recur. It was rewritten in September 2026 to reflect the Global Internal Audit Standards, the Basel III standardized approach that now governs capital in most jurisdictions, and the IIA’s Topical Requirements that touch operational risk. It sits in the site’s Risk Library beside the guides on credit, liquidity, interest rate, and compliance risk.

In this guide

Definition and boundaries

The Basel Committee’s definition has four causes (processes, people, systems, external events) and a set of boundaries that matter in practice. Legal risk is inside the definition: fines, penalties, damages, and settlements arising from the four causes are operational losses, which is why regulatory penalties for conduct failures sit in the operational risk loss database. Strategic risk and reputational risk are outside it: a bad acquisition is a strategic loss, and the customers who leave after an outage are a reputational consequence, even though the outage itself is an operational event. The boundary decisions matter because they decide what goes into the loss database and therefore into capital and into the board’s picture of the risk.

Operational risk also overlaps with every other risk stripe at the boundaries, and the overlaps are where losses are misclassified. A credit loss caused by a failed collateral process is an operational loss with a credit consequence, and Basel’s rules treat it as operational for management purposes and credit for capital; a market loss caused by an unauthorized trade is operational; a liquidity event caused by a payment system outage is operational. The OCC risk categories primer shows how the US regulators draw the stripes, and the audit question at each boundary is whether the institution has a rule for classifying boundary events and applies it consistently.

The Basel event-type taxonomy, with examples

The seven level-one event types are the classification most institutions use for loss data, RCSAs, and scenario analysis, and each has level-two categories beneath it. The table gives the definition, three concrete examples, and the control family that addresses each type, because a taxonomy is only useful if it points at controls.

Event typeDefinitionExamplesPrimary control families
1. Internal fraudLosses from acts intended to defraud, misappropriate property, or circumvent regulations or policy, involving at least one internal partyA settlement clerk diverting depot cash; a trader concealing positions; an employee creating a fictitious vendorSegregation of duties, access controls, reconciliations by independent staff, whistleblowing; see the fraud risk management guide
2. External fraudLosses from acts by a third party intended to defraud, misappropriate property, or circumvent the lawBusiness email compromise redirecting a supplier payment; card fraud; check kiting; first-party loan application fraudPayment verification and call-backs, customer authentication, transaction monitoring, positive pay
3. Employment practices and workplace safetyLosses from acts inconsistent with employment, health, or safety laws or agreements, and from discrimination or diversity eventsA wrongful-termination settlement; a warehouse injury claim; a class action on overtime classificationHR policy and training, safety programs, employment counsel review, incident reporting
4. Clients, products, and business practicesLosses from unintentional or negligent failure to meet a professional obligation to clients, or from the nature or design of a productFair-lending penalties; mis-sold products; fiduciary breaches; antitrust settlements; a product with an undisclosed feeCompliance management, product approval, complaints analysis, suitability controls, sales monitoring
5. Damage to physical assetsLosses from damage to physical assets from natural disasters or other eventsFlood damage to a data center; fire at a depot; vandalism of branchesInsurance, business continuity, facilities management, site selection
6. Business disruption and system failuresLosses from disruption of business or system failuresA core banking outage on a payroll day; a failed system migration; a cloud region failure; a telecom outage taking down call centersChange management, capacity and availability management, resilience testing, incident management; see the cybersecurity audit guide for the cyber-caused subset
7. Execution, delivery, and process managementLosses from failed transaction processing or process management, and from relations with trade counterparties and vendorsA wire sent to the wrong beneficiary; a collateral call missed; a reporting error to a regulator; a vendor failure that stops a process; a data-entry error that misprices a loan portfolioReconciliations, maker-checker, automated validations, vendor management; the Third-Party Topical Requirement governs the vendor part

Type 7 is where most events land by count and type 4 is where most losses land by value in banks, which is the pattern the industry loss consortia have reported for two decades: many small processing errors and a few very large conduct settlements. That distribution is why an operational risk framework needs two different instruments, a high-frequency process for the small events (reconciliations, incident logs, RCSAs) and a scenario process for the rare large ones, and why an audit that tests only the first has missed the losses that matter.

The operational risk management framework: components, artifacts, and what audit tests

An operational risk management framework (ORMF) is the set of components through which an institution identifies, assesses, monitors, controls, and reports operational risk. The Basel Committee’s Principles for the Sound Management of Operational Risk (revised 2021) set eleven principles; the components below are how those principles show up as things a person can inspect. The last column is the audit test, because a framework is audited through its artifacts and their use, not through its policy document.

ComponentPurposeOwnerArtifactWhat audit tests
Governance and policyBoard-approved framework, roles, and escalation; the ORM function’s mandate and independenceBoard; chief risk officerORM policy; committee charters and minutes; organization chartPolicy approval and currency; committee attendance and decisions; whether the ORM function reports outside the business lines it oversees
Risk taxonomyA common language for events, causes, and controls across the institutionORM functionTaxonomy document with level-one and level-two categories and boundary rulesConsistent use across RCSA, loss data, scenarios, and issues; boundary rules applied to a sample of events
Risk and control self-assessmentBusiness units identify their risks, rate them, and assess their controls; the first line’s own viewBusiness units; facilitated and challenged by ORMRCSA results by unit; challenge log; action plansCoverage of the unit’s processes; whether ratings changed after known events; evidence of second-line challenge; action plans tracked; see the RCSA process guide
Key risk indicatorsForward-looking metrics with thresholds that trigger action before a lossBusiness units; aggregated by ORMKRI inventory with definitions, thresholds, owners, and breach historyIndicator relevance to the top risks; threshold rationale; breaches actioned; indicators that never breach
Internal loss dataComplete, accurate record of operational loss events above a threshold, with causes and recoveriesBusiness units report; ORM validatesLoss database; event reports; reconciliation to the general ledgerCompleteness against the ledger (loss accounts, legal reserves, write-offs); timeliness; classification; near-miss capture; root-cause quality
External loss data and scenario analysisAssess exposure to rare, severe events using industry data and structured expert judgmentORM with business expertsScenario inventory with frequency and severity estimates and the workshop recordScenario coverage of the taxonomy; documented rationale for estimates; use in capital and appetite; refresh after events
Risk appetite and limitsBoard statement of the operational risk the institution will accept, cascaded to measurable limitsBoard; CROAppetite statement with operational risk metrics and limits; breach reportsMetrics actually measurable; breaches escalated and decided; see the risk appetite guide
Issue and action managementFindings from all sources tracked to closure with validationORM; internal audit for its ownIssue log with aging and validation statusSingle log or reconciled logs; aging measured from original dates; validation evidence; see the issue log template
Change and new-product risk assessmentOperational risk assessed before new products, systems, outsourcing, and reorganizations go liveBusiness units; ORM reviewNew-product approval records; change risk assessmentsSample of changes traced to an assessment; assessments completed before go-live; conditions tracked
Capital and reportingOperational risk capital calculated and reported; board and committee reporting on the risk profileFinance and ORMCapital calculation workpapers; board risk reportsData lineage for the calculation; reporting completeness (losses, KRIs, issues, scenarios, appetite); decisions taken on reports

Key risk indicators: twelve that work, with thresholds

A key risk indicator is a metric that moves before the loss does. Most KRI dashboards fail that test: they report losses that already happened (which are key performance indicators of the control environment, useful but late), or they report activity (training completion, RCSA completion) that measures the framework rather than the risk. The twelve below are leading indicators with a documented link to an event type, a starting threshold, and the rationale for it. Thresholds are calibrated to the institution’s own history, which is the second-line’s job; audit’s job is to ask for the calibration and to check that a breach produced an action. The site’s KRI reference table lists a much larger set by process.

#IndicatorEvent type it leadsStarting threshold (amber / red)Rationale
1Unreconciled items older than 30 days in cash, nostro, and suspense accounts, count and value7 Execution; 1 Internal fraudAny item over $25,000 / any over 60 daysAged differences are where processing failures and concealment live
2Payment repairs and returns as a percentage of payments7 Execution1.5 percent / 3 percentRising repair rates precede misdirected payments and fee losses
3Staff turnover in operations and control functions, trailing 12 months7 Execution; 1 Internal fraud15 percent / 25 percentKnowledge loss and untrained backups drive processing errors and segregation breaks
4Open high-rated audit and regulatory issues past original due dateAllAny over 90 days / any over 180 daysKnown control gaps that persist are the most predictable losses
5Privileged accounts without a named owner or with activity outside change windows6 Systems; 1 Internal fraudAny / more than 5Unowned privileged access precedes both outages and fraud
6Emergency changes as a percentage of all changes6 Systems10 percent / 20 percentBypassed change control is the leading cause of self-inflicted outages
7Critical vendors without a current risk assessment or with expired contracts7 Execution; 6 SystemsAny / more than 3Vendor failures arrive through the vendors nobody reassessed
8Complaints per 1,000 accounts, and the share escalated to a regulator4 Clients and productsTrend up 20 percent quarter on quarter / any regulator escalation clusterConduct losses are preceded by complaint patterns the institution had
9Manual journal entries as a share of all entries, and self-approved journals7 Execution; 1 Internal fraudTrend up / any self-approvedManual intervention rising means automated controls are being bypassed
10Overdue mandatory training and attestations in control-critical roles3 Employment; 4 Clients5 percent / 10 percentWeak on its own; strong when combined with turnover and complaints for the same unit
11Business continuity tests overdue or failed for critical processes5 Physical; 6 SystemsAny overdue / any failed without remediationAn untested plan is a plan that will fail on the day
12Near-miss events reported per unit, per quarterAllZero reported / declining while losses riseInverted: a unit that reports no near misses is not looking; the threshold catches silence

Two audit tests apply to every KRI program. First, take the last four quarters of breaches and trace each to a recorded decision: an action, an accepted risk, or a re-calibration with a reason. A breach with no decision means the indicator is reported and not used. Second, take the last four quarters of loss events above the reporting threshold and ask which indicator should have moved beforehand; where none did, either the indicator set has a gap or the thresholds are set where nothing ever breaches. Indicator 12 is the one to insist on, because the units with the best-looking dashboards are often the ones reporting nothing.

Measurement: frequency, severity, loss data, and capital

Operational losses are measured on two dimensions, how often events occur (frequency) and how large they are when they do (severity), and the two behave differently. Frequency is reasonably stable and predictable from history: a bank with 40,000 daily payments will have a steady rate of misdirected ones. Severity is heavy-tailed: most events cost hundreds or thousands, and a very small number cost hundreds of millions, and the tail events are not predictable from the body of the distribution. That shape is why an internal loss database, which mostly contains body events, cannot on its own tell an institution about its tail, and why scenario analysis and external loss data exist. The quantitative models that combined the two distributions into a capital number (the loss distribution approach) were permitted under Basel II’s advanced measurement approach; Basel III’s final standards replaced them with a single standardized approach for regulatory capital; the EU applied it from January 2025, while in the United States the agencies’ Basel III endgame rulemaking has been re-proposed and, as of September 2026, has not taken effect, so US institutions remain on the existing capital rules until it does. Under it, capital is a function of the institution’s size (the business indicator, a proxy built from income statement components) multiplied by an internal loss multiplier that scales with the institution’s own ten-year loss history, where national supervisors have not set the multiplier to one. The practical consequence for internal audit is that loss data completeness and accuracy became a capital input with a ten-year memory, and the audit of the loss database moved from a framework question to a capital-adequacy question.

Loss data has five quality attributes to test. Completeness: every event above the threshold is captured, which is tested by reconciling the database to the ledger accounts where losses land (operational loss expense, legal reserves and settlements, write-offs, customer remediation, fraud losses net of recoveries) and by searching incident, complaint, and legal logs for events the database lacks. Accuracy: gross loss, recoveries, and dates are right, tested to source documents. Timeliness: events are recorded within the policy window, tested by comparing event date to record date. Classification: event type and business line are consistent with the taxonomy, tested on a sample and on boundary events specifically. Root cause: the cause field says something a control owner can act on, which is a judgment tested by reading a sample; “human error” is not a root cause.

Worked example: what a small loss dataset tells you

Lakeshore Bancorp is a $9 billion regional bank with 60 branches, a mortgage company, and a wealth unit. Its loss database for the twelve months to 30 June held 214 events above the $5,000 reporting threshold, gross $4.62 million, recoveries $0.71 million. The table is the dataset summarized the way an operational risk committee should see it and an auditor should read it.

Event typeEventsGross lossLargest single eventMedian eventRecoveriesReading
1. Internal fraud3$186,000$142,000 (teller cash over 14 months)$22,000$61,000Low count, long duration: detection is slow; test dual control and surprise counts
2. External fraud71$1,240,000$310,000 (business email compromise, wire)$6,800$390,000Highest count; one large wire event is 25 percent of the type; call-back control failed once
3. Employment practices4$205,000$150,000 (settlement)$27,000$0Legal reserve reconciliation found one settlement not in the database
4. Clients and products9$1,910,000$1,400,000 (fee remediation on an overdraft program)$31,000$041 percent of total loss from one conduct event; complaints indicator had trended up for three quarters
5. Physical assets6$98,000$52,000 (branch flood)$11,000$74,000Insured; recoveries high; low residual
6. Business disruption11$312,000$180,000 (core outage, 6 hours, payroll day)$14,000$0Emergency-change indicator was red the month before the outage; nobody escalated
7. Execution and process110$669,000$96,000 (misdirected wire, unrecovered)$4,200$185,000Half of all events by count; 14 percent of loss; body of the distribution
Total214$4,620,000$710,000Two events are 37 percent of gross loss

The reading column is the analysis, and four things in it are the audit findings that came out of the engagement. The database was incomplete: reconciling to the legal reserve account found a $150,000 employment settlement that legal had booked and never reported, which is a completeness failure in a capital input. The conduct event had a leading indicator that was reported for three quarters and produced no decision, which is a KRI governance failure. The core outage followed a red emergency-change indicator, the same failure. And the median type-7 event of $4,200 sat below the $5,000 threshold, which means the database understates frequency for the most common event type and the incident log, not the loss database, holds the real picture of processing quality; the recommendation was a $1,000 threshold for type 7 with a lighter reporting form. None of those findings needed a model. They needed the ledger, the KRI history, and the incident log put next to the loss database.

Scenario analysis that produces defensible numbers

Scenario analysis is the framework’s instrument for the tail, and it fails in two predictable ways: estimates anchored on the last event the participants remember, and workshops where the most senior person’s number wins. A defensible process fixes both. Each scenario is written as a specific narrative with a cause, a control failure, and a consequence (“a business email compromise redirects three supplier payments over two weeks before the call-back exception is noticed”), not a category label. Participants estimate frequency and severity independently before the discussion, in writing, so the range is visible before anchoring sets in. Severity is estimated at two points, a typical-bad case (roughly one in ten years) and a severe case (roughly one in a hundred), each with the assumptions that produce it: how many payments, over how many days, at what average size, with what recovery. External loss data for the same event type is put on the table as a reference, not as the answer. And the record shows who said what and why, so that when the estimate is used in capital, appetite, or an insurance limit decision, the reasoning can be re-read. Audit’s test of a scenario program is the record: whether the estimates trace to stated assumptions, whether they changed after a real event of that type, and whether any decision anywhere in the institution was different because of them.

Who does what: the three lines in operational risk

Operational risk is the stripe where the three lines are most often confused, because the second-line function runs the framework and the first line runs the controls, and both call their work “risk management.” The first line, the business units and their embedded risk coordinators, owns the risks and the controls: it performs the RCSA, reports losses and near misses, maintains its KRIs, and fixes what breaks. The second line, the operational risk management function reporting to the CRO, owns the framework: it sets the taxonomy and methodology, challenges the first line’s assessments, aggregates and reports, runs scenario analysis, and monitors appetite. It does not own the controls, and an ORM function that “owns” the reconciliation control has crossed into the first line. The third line, internal audit, gives assurance over both: whether the first line’s controls work, and whether the second line’s framework is designed and operating so that the board’s picture of operational risk is true. The site’s page on the operational risk second line carries the RCM for that oversight role, and the guide to internal audit’s role in ERM sets out the roles audit must not take.

The common failure at the boundary is the second line doing the first line’s assessment for it, because the business will not, and then challenging its own work. Audit’s test is simple: pick three RCSAs and ask who typed the ratings. Where the ORM function did, the RCSA is a second-line opinion with a first-line label, and the first line’s ownership of its risks, which is the point of the exercise, is fiction.

Auditing the framework: a work program

An audit of the operational risk management framework is an audit of the second line and of the first line’s use of the framework, and it is distinct from the process audits (payments, lending operations, trade processing) that test operational controls directly. The program below is sized for a mid-size institution at 350 to 450 hours; it follows the components table and the work program conventions the site uses. Sample sizes follow the 25/40/60 guide unless the population is small enough to test in full.

#AreaObjectiveProcedureEvidenceHours
1GovernanceThe framework is approved, current, and overseenInspect the ORM policy and its approval; read four quarters of risk committee minutes for operational risk decisions; confirm the ORM head’s reporting linePolicy; minutes; org chart16
2GovernanceThe ORM function is independent of the business it challengesInspect role descriptions; interview the head of ORM and two business heads on how challenge works; test three challenge recordsRole descriptions; challenge log12
3TaxonomyOne taxonomy is used consistentlyCompare category use across RCSA, loss data, scenarios, and the issue log; test 25 boundary events against the boundary rulesTaxonomy; extracts from each system16
4RCSAAssessments cover the units’ processes and reflect realityMap RCSA scope to the process inventory; for 6 units, compare ratings before and after known loss events; test 25 action plans to closure evidenceRCSA results; loss events; action plans40
5RCSASecond-line challenge is realFor 25 RCSA lines, inspect the challenge record and whether the rating changed; interview three risk coordinators on who typed the ratingsChallenge log; interviews16
6KRIsIndicators are leading, thresholded, and acted onInspect the KRI inventory for definition, threshold rationale, and owner; trace all breaches in four quarters to a decision; map the year’s losses to indicators that should have movedKRI inventory; breach reports; loss data32
7Loss dataThe database is completeReconcile the year’s events to operational loss, legal reserve, write-off, and remediation accounts; search incident, complaint, and legal logs for unreported events above thresholdLedger extracts; logs; database40
8Loss dataEvents are accurate, timely, classified, and root-causedFor 40 events, test gross loss and recoveries to source, event-to-record days against policy, event type against the taxonomy, and the quality of the cause fieldEvent reports; source documents32
9ScenariosScenario analysis covers the tail and informs decisionsMap scenarios to event types and top risks; inspect workshop records for six scenarios; trace scenario outputs to appetite, capital, and insurance decisionsScenario inventory; workshop records; decision papers24
10AppetiteOperational risk appetite is measurable and enforcedInspect appetite metrics for measurability; test all breaches in four quarters to escalation and decisionAppetite statement; breach reports; minutes16
11IssuesIssues from all sources are tracked, aged honestly, and validatedReconcile ORM’s issue log to audit’s and to regulatory findings; test 25 closures for validation evidence; recompute aging from original datesIssue logs; closure evidence24
12ChangeNew products, systems, and vendors are risk-assessed before go-liveSelect 25 changes from project, product, and procurement records; trace each to an assessment dated before go-live; test conditions to closureChange records; assessments32
13CapitalLoss data feeding the capital calculation is controlledWalk the data lineage from the database to the calculation; test the ten-year data set for the completeness attributes; inspect the reconciliation between the twoCalculation workpapers; lineage24
14ReportingBoard and committee reporting is complete and produces decisionsInspect four quarterly reports for losses, KRIs, scenarios, issues, and appetite; trace three reported items to a decision in the minutesReports; minutes16
15Use testThe framework changes what the business doesFor three business units, trace one RCSA action, one KRI breach, and one loss event to a control change, a budget decision, or a product decisionUnit records; interviews24
Planning, reporting, review60

Procedure 15 is the one that distinguishes a framework audit from a compliance check. A framework can have every component on the table and change nothing; the use test asks whether a rating, a breach, or a loss ever altered a decision, and an institution where the honest answer is no has a reporting framework, not a risk management one. The findings from such an audit follow the severity scale like any other, with the loss-data completeness finding rated against its capital consequence.

Common failures in operational risk frameworks

FailureWhat it looks likeWhy it mattersFix
The green RCSAEvery unit rates every risk low and every control effective, year after year, including the year of the lossThe board’s picture of the risk is the business’s picture of itselfSecond-line challenge with teeth; ratings compared against loss and issue history; the ORM function reports the comparison
Loss data that stops at the thresholdA $5,000 threshold and a median type-7 event of $4,200The most common failures are invisible; capital inputs are incompleteTiered thresholds by event type; incident log linked to the database; near-miss capture
KRIs that never breachThresholds set at the 99th percentile of history; dashboards all greenIndicators reported, not usedCalibrate to trigger a few times a year; retire indicators with no breach in eight quarters; test breaches to decisions
Scenarios as a capital exerciseScenarios refreshed annually for the model and never read by the businessThe tail is estimated and not managedBusiness owners in the workshop; scenario outputs traced to insurance, appetite, and control decisions
Second line doing the first line’s jobORM types the RCSA, chases the losses, and owns the action plansOwnership is fiction; the second line challenges its own workNamed first-line risk coordinators; ORM restricted to facilitate and challenge
Three issue logsAudit, ORM, and compliance track findings separately; the committee sees three numbersNobody knows what is open; validation standards differOne log or a quarterly reconciliation; one aging rule from original dates
Change without assessmentNew products launched and vendors onboarded with the risk assessment done afterward, if at allThe framework sees the risk after the lossGo-live gated on assessment; audit tests the gate
Reports without decisionsForty-page quarterly packs; minutes that say “noted”The framework informs and nobody actsException-based reporting with a decision requested on each exception

Adapting: non-financial companies, small institutions, and the topical requirements

Outside financial services the vocabulary changes and the substance does not. A manufacturer or distributor has no operational risk capital and no Basel taxonomy, but it has processing failures, fraud, outages, vendor dependencies, and conduct exposure, and its enterprise risk register is where they live; the risk register guide and the risk-control matrix template carry the same logic without the regulatory layer. MidState Beverage, the site’s running example, has no ORM function; its route cash losses, depot inventory shrink, and handheld sync failures are operational risk by any definition, and the work program above, stripped of procedures 9 and 13, is the audit of how its management identifies and controls them. Small financial institutions run the framework with one person and a spreadsheet, and the audit tests are the same with smaller samples; what a small institution cannot skip is loss-data completeness, because the ledger reconciliation is a morning’s work and the regulator will ask.

Three of the IIA’s Topical Requirements sit inside operational risk and change what “must be covered” means: the Cybersecurity requirement (effective February 2026) governs event type 6’s cyber subset, the Third-Party requirement (effective September 2026) governs the vendor part of type 7, and the Organizational Resilience requirement (effective April 2027) will govern types 5 and 6 together. The topical requirements guide explains the conformance obligations; an operational risk framework audit is where a function shows it has considered all three. Model risk, which many institutions classify as operational, has its own guide on model risk audit, and the wider governance picture is in the GRC framework guide. Every guide on the site is indexed at All Guides.

Related guides

Comments

Leave a Reply

Discover more from internalauditguide.com

Subscribe now to keep reading and get access to the full archive.

Continue reading