A bank’s allowance for credit losses, its interest-rate risk position, its AML alert queue and most of its consumer credit decisions come out of models, and every one of those numbers reaches the board to two decimal places with no error bar attached. When a model is wrong, the loss rarely announces itself: an allowance runs a few million dollars light for five quarters, a monitoring threshold gets tuned until the alert backlog is manageable, a pricing spreadsheet keeps a two-year-old funding curve. Validation teams exist to catch this, and internal audit exists to know whether validation is working. Most model risk audits fail that second job in one of two ways: they re-validate one model for 600 hours and say nothing about the other hundred, or they tick through the policy and never look at a number.
This guide gives you a model inventory and tiering table with scoring criteria, a 20-test model risk management (MRM) audit work program, a checklist for reading a validation report, a deep-dive procedure with a worked CECL example that ends in a backtest and a rated finding, and a table of the findings that recur in almost every MRM audit. It was rewritten in September 2026 to reflect the Global Internal Audit Standards and the revised interagency model risk guidance that replaced SR 11-7 on 17 April 2026. It pairs with the guide to auditing AI and algorithms, which covers fairness and generative systems, and with the operational risk guide, where model failure sits as a loss event type.
In this guide
- What SR 11-7 built, and what the April 2026 revision changed
- Model inventory and tiering: the table the audit is built on
- The MRM audit work program: seven areas, 20 tests
- How to read a validation report
- When internal audit deep-dives a model: a worked CECL example
- Spreadsheets, EUCs and machine-learning models
- Common findings, and how the audit itself goes wrong
- Adapting the program: smaller banks, insurers, corporates and the committee
What SR 11-7 built, and what the April 2026 revision changed
For fifteen years the reference document was the Federal Reserve’s SR 11-7 and its OCC twin, Bulletin 2011-12, “Supervisory Guidance on Model Risk Management,” issued 4 April 2011 and adopted by the FDIC in 2017. It defined a model as a quantitative method, system or approach that applies statistical, economic, financial or mathematical theories, techniques and assumptions to process input data into quantitative estimates, expert-judgment approaches included whenever the output was quantitative. It defined model risk as the potential for adverse consequences from decisions based on incorrect or misused model outputs, arising from fundamental errors or from incorrect use. It named the three elements every validation report is still organized around (conceptual soundness, ongoing monitoring, outcomes analysis) and set validation’s standard of proof, “effective challenge,” which depends on incentives, competence and influence rather than on a signature.
On 17 April 2026 the OCC (Bulletin 2026-13), the Federal Reserve (SR 26-2) and the FDIC jointly issued revised guidance that supersedes SR 11-7, the 2021 interagency statement on MRM for BSA/AML systems, and the OCC’s 2021 Comptroller’s Handbook booklet on model risk. The three validation elements, effective challenge, the inventory, and the governance and vendor chapters all survived; what changed is scope and tone. A model is now a “complex” quantitative method applying statistical, economic or financial theories; simple arithmetic such as that found within spreadsheets and deterministic rule-based processes are expressly excluded, and generative and agentic AI are out of scope as “novel and rapidly evolving.” The guidance is “expected to be most relevant to banking organizations with over $30 billion in total assets,” it “does not set forth enforceable standards or prescriptive requirements,” and non-compliance “will not result in supervisory criticism.” The mandated annual review is gone; validation frequency now varies with purpose, methodology, change, data limitations and practical constraints. And materiality is defined: model purpose, together with model exposure, determines it.
Internal audit’s paragraph is shorter but says the same thing. Internal audit “would generally not duplicate model risk management activities such as model development or validation”; its role “is generally to evaluate whether the model risk management practices are rigorous and effective and whether related policies are implemented accordingly.” That sentence is the whole engagement: audit the practices, test the policy against what happened, and go into an individual model only when a decision rule says the practices cannot be relied on for it. The evaluation criteria under Standard 13.4 of the Global Internal Audit Standards are therefore the bank’s own policy, and model risk itself is not one of the OCC’s eight risk categories but a driver that runs through them, as the OCC risk categories primer explains. The table lists the changes that alter how you scope and write the audit.
| Topic | SR 11-7 / OCC 2011-12 (2011) | SR 26-2 / OCC 2026-13 (17 April 2026) | What it changes in the audit |
|---|---|---|---|
| Definition of a model | Any quantitative method applying statistical, economic, financial or mathematical theory; expert-judgment inputs included if the output is quantitative | A “complex” method applying statistical, economic or financial theory; spreadsheet arithmetic and deterministic rules excluded | Inventories shrink; test each removal’s rationale and that the item landed on the EUC register with controls |
| Who it applies to | All supervised banks in proportion to size; FDIC’s 2017 adoption drew a $1 billion line | “Most relevant” above $30 billion; may apply below where model exposure is significant | Below $30 billion the criteria are the bank’s own policy, not the guidance; say so in the planning memo and report |
| Enforceability | Guidance treated by examiners as the expectation | No enforceable standards; non-compliance draws no supervisory criticism | Write findings against the policy the board adopted, never against the guidance as if it were a rule |
| Drivers of model risk | Complexity, uncertainty about inputs and assumptions, extent of use, potential impact | Inherent risk, exposure, purpose and use; purpose plus exposure sets materiality; immaterial models may get basic monitoring only | Map tiering criteria to the four drivers; test that every “immaterial” label is documented and approved, since it removes validation |
| Validation cadence | Periodic review of each model at least annually | No mandated cadence; frequency follows purpose, methodology, change, data and practical constraints | Test that policy sets a frequency per tier, that it is met, and that overdue models carry interim controls |
| Independence of validation | Validators independent of development; effective challenge from incentives, competence and influence | Objective experts with sufficient independence; quality depends on rigor rather than organizational structure | Test rigor directly: did the validator replicate results, raise findings, get overruled? The org chart no longer answers the question |
| BSA/AML systems | 2021 statement applied MRM principles flexibly | 2021 statement rescinded; AML systems meeting the definition are in scope like any other | Scenario tuning and threshold changes are model changes; scope them once, with the BSA/AML audit |
| Vendor models | Developmental evidence, validation, contingency plans | Understand design and performance, validate, monitor outcomes, document customizations; contingency language dropped | Keep testing exit and contingency under the third-party program, where the IIA’s Third-Party Topical Requirement (effective 15 September 2026) expects them |
Model inventory and tiering: the table the audit is built on
The inventory is the population for every other test in the audit and the thing most likely to be wrong, so it is tested before it is used. The revised guidance asks for “sufficient information to understand model risks, so as to support effective model risk management at the individual and aggregate levels”: one row per model with owner, developer, validator, purpose, methodology, tier and score, upstream and downstream models, version, status, validation dates and result, open findings, monitoring status and known limitations. The worked example throughout is Sable Creek Bancorp, a privately held $9.6 billion bank holding company with $7.1 billion of loans and an allowance for credit losses (ACL) of $88.0 million (1.24 percent of loans) at 31 December 2025. Its inventory holds 112 models: 11 in Tier 1, 34 in Tier 2 and 67 in Tier 3, 40 of them spreadsheets. MRM is a four-person team under the CRO; Tier 1 validations are outsourced, the rest are internal, and a Model Risk Committee (MRC) chaired by the CRO meets quarterly. The policy scores four criteria that map to the four drivers in the revised guidance, each 1 to 3 against the thresholds below; the sum sets the tier, and published thresholds make a tier re-performable, which is exactly the test in the work program.
| Criterion | Score 1 | Score 2 | Score 3 |
|---|---|---|---|
| Exposure (how much rides on the output) | Under $10 million of balances, earnings or capital, or decisions reversible within a quarter | $10 million to $250 million, or decisions affecting one business line | Over $250 million, a financial statement estimate, a capital or liquidity measure, or automated decisions on customers |
| Purpose (what the output is used for) | Internal analysis, planning, advisory use with expert review | Management reporting, pricing, limit setting, staffing | Financial reporting (allowance, valuations), regulatory reporting, consumer credit decisions, BSA/AML alerting |
| Inherent risk (how likely it is to be wrong) | Transparent formula, stable data, few assumptions | Statistical estimation (regression, time series), documented assumptions, moderate data limitations | Complex or opaque methods (simulation, machine learning, vendor engine), judgment-heavy inputs, proxy data, macro forecasts |
| Use (how it is applied) | Single team, every output reviewed by an expert before use | Several users, periodic review, occasional overrides | Enterprise-wide or straight-through automated decisions, outputs feed other models, overrides frequent |
Scores of 10 to 12 make a model Tier 1, 7 to 9 Tier 2, and 4 to 6 Tier 3, and each tier carries a validation scope and frequency in the policy, which is what replaces the annual review the 2011 guidance mandated. At Sable Creek, Tier 1 means full external validation of all three elements with replication of key results before first use, after material change and every 24 months, plus quarterly outcomes analysis against written thresholds with breaches escalated to the MRC the following quarter; Tier 2 means full internal validation every 36 months with semi-annual monitoring; Tier 3 means targeted validation of conceptual soundness and implementation before use and on material change, with an annual owner attestation. Tools reclassified as non-models leave MRM for the EUC standard: access, version and change controls, input reconciliation and formula integrity, attested annually.
In the extract below, two rows are findings before fieldwork starts: the commercial risk-rating scorecard is nine months past its validation date, and the marketing propensity model has never been validated. The RAROC calculator is the 2026 re-scoping question: MRM proposed moving it and 30 other Tier 3 spreadsheets to the EUC register as non-models, defensible only if that register exists and has controls, which at Sable Creek it did not.
| Model | Purpose and method | Owner | Exposure / Purpose / Inherent / Use | Score, tier | Status at 30 June 2026 |
|---|---|---|---|---|---|
| CECL allowance model v3.4 | Lifetime expected loss on $7.1 billion of loans; vendor platform, discounted cash flow with PD/LGD term structures for commercial segments, vintage loss rates for consumer; nine qualitative factors totaling $8.4 million | Chief Credit Officer | 3 / 3 / 3 / 3 | 12, Tier 1 | Validated Q3 2024 externally, “approved with conditions” (2 high, 4 medium, 5 low); annual review memo December 2025 |
| ALM earnings and EVE simulation | Interest rate risk in the banking book across 12 rate scenarios; vendor engine, bank-set prepayment and deposit assumptions | Treasurer | 3 / 3 / 2 / 2 | 10, Tier 1 | Validated Q1 2025, approved; monitoring current |
| Commercial risk-rating scorecard | PD and LGD grades on 3,900 relationships; in-house logistic regression on 11 ratios plus qualitative inputs | Credit Administration | 3 / 2 / 2 / 3 | 10, Tier 1 | Last validated Q4 2023; nine months overdue, no interim control |
| Consumer and indirect auto scorecard | Bureau score plus custom cutoffs; 71 percent of applications decided automatically | Head of Consumer Lending | 3 / 3 / 2 / 3 | 11, Tier 1 | Validated Q2 2025; override rate 6.8 percent against a 5 percent limit |
| Marketing propensity model | Vendor-built gradient-boosted classifier on 1.4 million customer records; selects campaign targets | Head of Marketing | 1 / 1 / 3 / 2 | 7, Tier 2 | Never validated; no drift monitoring; found by the 2025 discovery scan |
| Loan pricing RAROC calculator | Spreadsheet applying a deterministic formula to funding curve, capital and expected loss inputs | Commercial Banking | 2 / 1 / 1 / 2 | 6, Tier 3 | Proposed non-model 2026; funding curve last refreshed August 2024 |
The inventory should reconcile to the enterprise risk register, where aggregate model risk is an entry with a rating, an owner and indicators (overdue validations, open high findings, monitoring breaches); when the ERM report and the MRM report disagree for the same quarter, one of them is being produced from memory.
The MRM audit work program: seven areas, 20 tests
Sable Creek’s 2026 MRM audit was budgeted at 380 hours: an audit manager, a senior with an FRM and two years in credit risk, and 60 hours of a co-sourced quantitative specialist for the deep-dive, an arrangement the co-sourcing guide covers and one that Standard 3.1 on competency makes hard to avoid when nobody on the team has estimated a regression. Fieldwork ran in June 2026, two months after the agencies replaced SR 11-7; the board reaffirmed the policy with updated references on 24 June, so the criteria are policy v4.2. The planning-memo language follows, in the shape the audit work program guide recommends.
Objective. To assess whether Sable Creek’s model risk management practices are rigorous and effective and whether the Model Risk Management Policy (v4.2, approved 24 June 2026) is implemented across the 112 models, and to test whether the CECL allowance model (v3.4) performed as intended for the year ended 31 December 2025.
Scope. Governance and reporting; inventory and tiering; development, implementation, change and use controls; independent validation; monitoring and outcomes analysis; vendor models; and issue management, for 1 July 2025 to 30 June 2026. One deep-dive of the CECL model’s commercial real estate probability-of-default component, selected under the decision rule at workpaper A-4.
Out of scope. Re-validation of any model. The vendor CECL platform’s discounted cash flow engine, validated externally in September 2024 and relied upon under Standard 9.5 (reliance assessment at workpaper A-7). Generative AI tools, covered by the Q4 2026 AI governance review.
Criteria. MRM Policy v4.2 and its Validation, Monitoring and EUC Standards; the revised interagency guidance (SR 26-2 / OCC Bulletin 2026-13) as adopted by that policy; ASC 326 and the Interagency Policy Statement on Allowances for Credit Losses (2020, revised 2023) for the deep-dive; FDICIA Part 363 for the ICFR implication.
Attribute samples follow the site’s 25/40/60 convention, full-population analytics are used wherever the data exists, and selections are made with mySampler with the seed recorded. Every test names its evidence, because the most common weakness in MRM workpapers is a conclusion resting on an interview.
| Ref | Area | Objective | Test | Evidence and sample |
|---|---|---|---|---|
| A1 | Governance | The MRC and board risk committee oversee model risk as the policy describes | Read four quarters of MRC minutes and board packs; trace three escalated items to a recorded decision; check quorum and agenda against the charter | MRC charter; minutes Q3 2025 to Q2 2026 |
| A2 | Governance | The policy is current and complete | Map policy v4.2 and its standards to the seven sections of the revised guidance; list gaps (no consequence for a monitoring breach; no EUC standard until June 2026) | Policy; mapping workpaper |
| A3 | Governance | Validation has the incentives, competence and influence to challenge | Review MRM staffing, qualifications and reporting line; scan the findings log for findings overridden by the business or MRC; confirm no validator validated a model they built; check engagement letters for scope limits | Org chart; CVs; findings log (63 items); engagement letters |
| B1 | Inventory | The inventory is complete | Build a candidate list from outside the inventory (application inventory, EUC register, contracts containing “model,” “score,” “analytics” or “forecast,” GL allowance and valuation accounts, change tickets citing parameters); reconcile and disposition every difference | Candidate list (137 items); nine unlisted; dispositions signed by MRM |
| B2 | Inventory | Records are accurate and tiering is re-performable | Sample 25 records and agree owner, tier, version, status, last validation date and result, and open findings to source documents; re-score all 11 Tier 1 and 15 Tier 2/3 models, investigate differences of 2 points or more, and confirm each downgrade since prior year has an MRC decision | 25 records; validation reports; scoring sheet; MRC minutes for four downgrades |
| B3 | Inventory | 2026 non-model reclassifications are justified and controlled | Review the rationale for each of 31 items proposed for removal; confirm each is on the EUC register with an owner and controls; test five | Reclassification memo; EUC register; five spreadsheets |
| C1 | Development and use | Documentation and implementation evidence let an unfamiliar reviewer understand and trust the model | For three Tier 1 and two Tier 2 models, assess purpose, data lineage, assumptions, alternatives, testing and limitations; obtain the go-live and last-version reconciliation of production to development output; re-perform one | Five development documents; UAT sign-offs |
| C2 | Development and use | Changes are classified, tested, approved and validated when material | List all 2025 changes from the platform audit log and repository (41); sample 15 for classification, testing, approval, and whether “material” triggered validation | Change log; 15 tickets; repository history |
| C3 | Development and use | Access to code and parameters is restricted | Obtain user lists for the CECL platform (14 users, three with parameter-edit rights) and the scorecard; check role, last access review and leavers; confirm developers cannot promote to production | User lists; access review; leavers list |
| C4 | Development and use | Overrides are limited and known limitations carry operating compensating controls | Compute override rates for the consumer scorecard and risk-rating model against policy limits and sample 25 for approval and rationale; for each Tier 1 model, pair every limitation in the last validation with its compensating control and test that it operated | Override logs; 25 overrides; limitations register |
| D1 | Validation | Coverage meets policy frequency | Compare last validation dates to tier frequency for all 112 models; for each overdue model (seven Tier 1/2), test for an interim control or MRC-approved extension | Inventory; validation schedule |
| D2 | Validation | Reports cover all three elements with numbers | Apply the checklist below to four reports (two external, two internal); record whether outcomes analysis states numbers against thresholds and whether key results were replicated | Four validation reports |
| D3 | Validation | Validation work can be relied on (Standard 9.5) | Assess objectivity, competence and evidence quality of the validation function; document the reliance decision and the tests reliance does not remove | Reliance assessment at A-7 |
| D4 | Validation | Findings are remediated and closed on evidence | From 63 findings, sample 20 closed; check closure evidence, validator re-test, MRC approval and aging against policy (high findings 90 days) | Findings log; 20 closure packages |
| E1 | Monitoring | Every Tier 1 and 2 model has a monitoring plan with meaningful thresholds and a consequence for breach | For the 45 Tier 1/2 models, confirm metrics, thresholds, frequency, owner and escalation path; assess the statistical basis of thresholds and what the standard requires on breach | 45 monitoring plans |
| E2 | Monitoring | Reports are accurate, acted upon, and backed by credible outcomes analysis | For eight models, obtain four quarters of reports; test the extract as IPE, recompute two metrics per model from source data, check breaches against the escalation path; confirm a backtest or benchmark for each Tier 1 model and re-perform one in full | 32 reports; IPE workpaper; MRC minutes; 11 backtests |
| F1 | Third-party models | The bank understands its vendor models and their customizations | For the five vendor models, confirm design documentation, vendor testing evidence and the bank’s own written understanding; extract bank-set parameters and verify each was justified, approved and inside validation scope | Vendor files; configuration extracts |
| F2 | Third-party models | Contracts and contingency support model risk management | Check contract rights to validation evidence and change notification; check the exit or contingency plan under the third-party program | Five contracts; TPRM files |
| G1 | Issues and reporting | Regulatory and prior audit findings are remediated on evidence | Review examination comments and any MRA on model risk and verify remediation evidence; validate closure of the four open findings from the 2025 audit | Exam reports; remediation tracker; issue log |
| G2 | Issues and reporting | Aggregate model risk reporting is accurate and reaches the board | Recompute the Q4 2025 and Q1 2026 MRM report counts (models by tier, overdue validations, findings by age, breaches) from the inventory and log; agree to the ERM report and risk register | Two MRM reports; ERM report; reconciliation |
Three tests carry most of the weight. B1 is the only real test of completeness, because asking MRM whether the inventory is complete asks the population to certify itself; at Sable Creek the scan found nine candidates, four of them models (the propensity model, a vendor deposit-pricing optimizer, an OREO valuation haircut and a branch-closure NPV tool). E2 is where IPE testing earns its place: a monitoring report is information produced by the entity, and recomputing a metric without testing the extract proves only that the arithmetic works on whatever went in. D3 is the reliance decision, and it deserves the rigor of any other audit evidence: who did the work, what qualified them, what they did, what they found.
How to read a validation report
A validation report is the best evidence in the audit and the easiest to misread, because a 90-page report with a green “approved” on the cover looks like assurance whether or not it contains any. The checklist gives, section by section, what a rigorous report contains, the red flag that means the section is decorative, and the document to verify it against; score each element present or absent so the reliance decision rests on a count. A validation that describes the backtest it would have run has tested design only, the distinction the design versus operating effectiveness guide draws for any control.
| Report section | What a rigorous one contains | Red flag | Verify against |
|---|---|---|---|
| Scope statement | Model name, version, components and uses covered, data as-of date, exclusions and why | No version number; “the model” without saying which build | Inventory record; change log since the as-of date |
| Independence and staffing | Names, roles, qualifications, reporting line; prior involvement in development disclosed | Validator also credited on the development document | Development document authors; org chart |
| Conceptual soundness | Theory, alternatives tested and rejected, variable selection evidence, assumptions with the output’s sensitivity to each | “Consistent with industry practice” with no alternative tested | Development document; sensitivity tables |
| Data | Lineage from source to model input, sample period, exclusions and their effect, reconciled counts and balances | No reconciliation; sample period ends before a known stress episode | Extract reconciliation; GL |
| Replication | Validator re-estimated key parameters or re-ran the model and reports the difference | Results “reviewed” rather than reproduced | Validator’s code or workbook |
| Implementation testing | Production output agreed to development output on a defined test set; parameter checks | Implementation “outside the scope of this validation” | UAT evidence; C1 workpaper |
| Outcomes analysis | Predicted against actual at the grain the model predicts, a statistical test, a threshold set in advance, a numeric result | “Performance is acceptable” with no table | Deep-dive re-performance; monitoring reports |
| Limitations and findings | Each limitation paired with a control and owner; findings with severity, action, owner and date; conditions of approval separated from recommendations | Limitations without controls; findings without dates; “approved” with high findings open | C4 limitations register; findings log; MRC minutes |
| Vendor addendum | Developmental evidence obtained, customizations tested, bank-side outcomes analysis | “Validated by the vendor” offered as the bank’s validation | F1 workpaper |
When internal audit deep-dives a model: a worked CECL example
The guidance says internal audit should not duplicate validation, and the temptation is strong because a model is more interesting than a policy. The discipline is a decision rule agreed with the CAE in planning: a model is deep-dived only when a trigger below fires, and the deep-dive is limited to the component and question the trigger raises. It re-performs one element of outcomes analysis, quantifies what the result means for the number management relies on, and traces how governance handled evidence it already had. Where D3 finds the validation function rigorous, one deep-dive a year is enough; where it does not, plan for three and say why, as Standard 13.2 on engagement risk assessment requires.
| Trigger | Threshold | Decision |
|---|---|---|
| Validation overdue | Tier 1 or 2 model past policy frequency, no MRC-approved extension, no interim control | Deep-dive the component with the largest exposure |
| Conditional approval with open conditions | Any high or medium condition open more than 180 days, or closed without validator re-test | Deep-dive the component the condition concerns |
| Monitoring breach without escalation | Threshold breached in two or more consecutive periods with no recorded MRC decision | Re-perform the breached metric from source data |
| Material change since validation | Methodology, data, vendor version or key parameter changed and classed non-material by the owner | Re-run before and after on a fixed test set |
| Financial reporting or regulatory output | Output feeds the allowance, a fair value, the call report, capital or liquidity reporting; no other trigger fired | Candidate; one a year on rotation, coordinated with the ICFR program |
| None of the above | Validation current, conditions closed on evidence, monitoring within thresholds | Rely on validation (Standard 9.5); test the practices only |
At Sable Creek the CECL model fired two triggers: finding M-3 from the 2024 validation had been closed in February 2025 on a management overlay without validator re-test, and the quarterly backtest of the commercial real estate PD component had breached its threshold for three consecutive quarters with no MRC decision on record.
- Fix the question. Model v3.4, component: one-year probability of default for non-owner-occupied commercial real estate (NOO CRE), as of 31 December 2024 with outcomes through 31 December 2025. Does the component predict defaults within tolerance, and if not, what does the miss do to the allowance?
- Obtain and test the inputs as IPE. Loan-level extract at 31 December 2024 (1,312 loans, $2.35 billion) reconciled by count and balance to the general ledger and the model input file; 25 loans traced to the loan system.
- Read the development document and the 2024 validation against the checklist. Limitations: a single macroeconomic driver (national unemployment), a four-quarter reasonable and supportable period, straight-line reversion over four quarters to the 2005–2024 average.
- Re-perform the outcomes analysis at the model’s own grain. Expected defaults by risk grade from the model’s one-year PDs against 2025 defaults (90 days past due, non-accrual or charge-off), with an exact binomial test and a traffic-light rule fixed in advance: green above p = 0.10, amber from 0.05 to 0.10, red at or below 0.05.
- Quantify the effect on the number management relies on. The owner re-ran the 31 December 2025 segment in the test environment with grade 6 to 8 PDs recalibrated as the validator had recommended; the audit computed an independent approximation to corroborate it.
- Compare with the compensating control. The office-concentration overlay sized against the gap the re-run showed.
- Trace the governance trail. Monitoring reports Q2 to Q4 2025, MRC minutes, the M-3 closure package, the December 2025 annual review memo.
- Write the finding at the control level with the dollar effect as a range. Facts agreed with the owner and the validator before drafting; the ICFR implication referred to the FDICIA program under the site’s control deficiency evaluation approach.
The backtest
The model assigns each NOO CRE loan a one-year PD by risk grade; grades 1 to 3 held no loans in the segment. The shortfall is not subtle: 34 defaults against 18.0 expected, a ratio of 1.89 and a z-score of 3.9, which puts the chance of that many defaults under correct PDs below one in ten thousand. The miss sits in grades 6 and 7, the watch and special-mention grades where office exposure ($672 million, 28 percent of the segment) is clustered.
| Risk grade | Loans at 31 Dec 2024 | Model one-year PD | Expected defaults | Observed 2025 | Observed rate | Actual to expected | Binomial p (one-sided) | Result |
|---|---|---|---|---|---|---|---|---|
| 4 (pass) | 612 | 0.35% | 2.1 | 3 | 0.49% | 1.40 | 0.36 | Green |
| 5 (pass) | 438 | 0.80% | 3.5 | 5 | 1.14% | 1.43 | 0.27 | Green |
| 6 (watch) | 171 | 2.10% | 3.6 | 9 | 5.26% | 2.51 | 0.011 | Red |
| 7 (special mention) | 54 | 6.00% | 3.2 | 8 | 14.81% | 2.47 | 0.015 | Red |
| 8 (substandard) | 37 | 15.00% | 5.6 | 9 | 24.32% | 1.62 | 0.092 | Amber |
| Segment | 1,312 | 1.37% (count-weighted) | 18.0 | 34 | 2.59% | 1.89 | Under 0.0001 (z = 3.9) | Red |
The monitoring reports had already seen this coming: the trailing four-quarter actual-to-expected ratio stood at 1.31 in Q2 2025, 1.58 in Q3 and 1.89 in Q4 against a band of 0.75 to 1.25. Each report labeled the metric “informational,” and the MRC minutes record it as received without discussion. The monitoring standard defined the threshold but not what a breach obliged anyone to do, the design gap that turned three amber quarters into nothing.
What the miss does to the allowance
Finding M-3 had recommended a CRE price or vacancy driver, or recalibration of grade-level PDs to the bank’s own defaults. The owner re-ran the 31 December 2025 segment ($2.40 billion) with grade 6 to 8 one-year PDs blended 50/50 with the trailing eight-quarter observed rates, grades 4 and 5 unchanged; the segment’s quantitative reserve moved from $27.5 million to $31.9 million, up $4.4 million. The audit’s approximation below applies the same blended PDs to first-year expected loss only, ignores discounting and reversion, and lands at $4.2 million: close enough to corroborate the platform result and simple enough to show a committee in one table.
| Risk grade | Balance at 31 Dec 2025 ($M) | Model one-year PD | Blended PD | Model LGD | First-year expected loss, model ($M) | First-year expected loss, blended ($M) | Difference ($M) |
|---|---|---|---|---|---|---|---|
| 4 | 1,180 | 0.35% | 0.35% (unchanged) | 25% | 1.03 | 1.03 | 0.00 |
| 5 | 760 | 0.80% | 0.80% (unchanged) | 25% | 1.52 | 1.52 | 0.00 |
| 6 | 290 | 2.10% | 3.68% | 30% | 1.83 | 3.20 | 1.38 |
| 7 | 105 | 6.00% | 10.41% | 35% | 2.20 | 3.82 | 1.62 |
| 8 | 65 | 15.00% | 19.66% | 40% | 3.90 | 5.11 | 1.21 |
| Segment | 2,400 | 1.35% (balance-weighted) | 1.86% (balance-weighted) | 26% (balance-weighted) | 10.48 | 14.69 | 4.21 |
Against that, the office-concentration overlay was $2.1 million, set according to the Q-factor memo “based on management judgment regarding office market conditions,” with no reference to the backtest, the validation finding or any calculation. It covered less than half of the gap the bank’s own data supported, and the MRC had accepted it as closure of M-3 without asking the validator to look. The December 2025 annual review memo, two pages, stated that no material changes had been made and concluded the 2024 validation remained current; it did not mention the monitoring results. Nothing in this chain required a quantitative specialist, only someone reading the monitoring report, the finding and the closure package side by side.
The finding
The finding is written against policy in the five Cs structure, with the dollar effect as a range and the rating justified under the site’s severity scale. It is a control finding first: the reason it is High is that three controls designed to catch a miscalibrated component all operated and none produced a decision.
Condition. The CECL model’s one-year probability of default for non-owner-occupied commercial real estate under-predicted 2025 defaults by 47 percent: 34 observed against 18.0 expected across 1,312 loans (p under 0.0001), with grades 6 and 7 individually failing at the 5 percent level. Quarterly monitoring showed the actual-to-expected ratio at 1.31, 1.58 and 1.89 for Q2 to Q4 2025 against a 0.75 to 1.25 band; no escalation or MRC decision was recorded. Validation finding M-3 (September 2024), recommending recalibration or an additional CRE driver, was closed by the MRC in February 2025 on a $2.1 million qualitative overlay with no quantitative support. Recalibrating grade 6 to 8 PDs as recommended raises the 31 December 2025 quantitative reserve by $4.4 million, $2.3 million more than the overlay.
Criteria. MRM Policy v4.2 section 5.3 (monitoring breaches escalated to the MRC the following quarter with a documented decision); section 6.4 (validation findings closed only on evidence reviewed by the validator); the Interagency Policy Statement on Allowances for Credit Losses (qualitative adjustments supported and documented).
Cause. The monitoring standard labels the backtest informational and sets no consequence for a breach; the MRC accepted an overlay as closure without validator re-test; the annual review relied on a “no material change” attestation and ignored monitoring results.
Consequence. The allowance at 31 December 2025 may be understated by $2.3 million to $4.4 million (2.6 to 5.0 percent of the ACL), below the external auditor’s $5.9 million overall materiality but above the $1.2 million reporting threshold agreed with the audit committee. The bank relied for five quarters on a component its own data showed to be miscalibrated, and the FDICIA Part 363 evaluation must weigh the control gap, not only the dollar effect.
Corrective action (agreed). Credit Administration will recalibrate grade 6 to 8 PDs and add a CRE price index driver for the 30 September 2026 allowance, validated before use (Chief Credit Officer, 30 September 2026). MRM will amend the monitoring standard so two consecutive breaches require a recorded MRC decision (Head of MRM, 31 August 2026). The MRC will reopen M-3 for validator re-test before closure (CRO, 31 July 2026). Rating: High.
The same procedure transfers to other model families with a change of metric. For the ALM model it is realized net interest income against the projection for the rate path that occurred, using the measures in the IRR, IRRBB and SIRR guide. For a VaR model it is the count of exceptions in 250 trading days against the Basel traffic-light zones (green at 0 to 4, yellow at 5 to 9, red at 10 or more for a 99 percent one-day VaR), covered in the VaR guide. For a liquidity model it is projected against actual deposit runoff in a stress month (liquidity risk guide); for an AML system, below-the-line testing of the alerts that tuned thresholds suppressed, inside the AML/KYC audit; for a risk-rating scorecard, the migration and default experience of each grade, a standing test in the credit risk guide.
Spreadsheets, EUCs and machine-learning models
The 2026 definition pushes spreadsheets out of the model inventory unless they carry statistical, economic or financial theory, and the honest reading is that most do not: a RAROC calculator applying a formula to inputs is arithmetic, while a deposit-decay regression in a workbook is a model whatever file it lives in. The audit risk is the transition. When Sable Creek’s MRM proposed reclassifying 31 Tier 3 spreadsheets, the memo named the items and the rationale but had nowhere to send them, because the bank had no EUC standard; 31 tools that had at least received an annual owner attestation were about to receive nothing. The audit’s position, which the MRC accepted, was that reclassification is fine on the day the EUC register exists with an owner, access and version control, a change test and an input reconciliation for each item, and not before. The RAROC calculator’s funding curve, last refreshed in August 2024 and pricing 2026 loans off a two-year-old cost of funds, made the point in the room.
Machine-learning models that are not generative sit inside the definition: a gradient-boosted classifier applies statistical theory to produce a quantitative estimate, and “complex” describes it exactly. The tests are the same seven areas with four additions. Data and feature governance: training data lineage, exclusion rules, and whether protected characteristics or their proxies (ZIP code, first name, device type) are features, which for consumer credit engages the Equal Credit Opportunity Act and Regulation B, including the duty to state specific principal reasons on an adverse action notice. Drift monitoring: population stability index on inputs and score distribution, with the working thresholds of 0.10 for a moderate shift and 0.25 for a significant one. Retraining governance: each retrain is a model change, and the policy must say when one needs validation before deployment. Explainability: whether the validator could reproduce the feature-importance ranking and whether users receive reason codes that mean something. Sable Creek’s marketing propensity model failed all four; an exposure score of 1 did not stop an inherent-risk score of 3 with no validation from being a finding on its own.
Generative and agentic systems are outside the guidance, which does not put them outside the audit universe. NIST AI RMF 1.0 (2023) supplies the Govern, Map, Measure and Manage functions, ISO/IEC 42001:2023 a certifiable management system, and for a bank with EU customers Regulation (EU) 2024/1689 lists creditworthiness assessment of natural persons and life and health insurance pricing among its Annex III high-risk uses; as of September 2026 those obligations apply from 2 December 2027, after the Digital Omnibus on AI amendments adopted in mid-2026 deferred the original 2 August 2026 date. The audit question is coverage: which program (MRM, AI governance, third-party risk, information security) owns each system, and whether any falls between them. The AI and algorithm audit guide carries the fairness tests, explainability procedures and generative-system controls this guide does not repeat.
Common findings, and how the audit itself goes wrong
The findings below account for most of what an MRM audit reports in a bank between $5 billion and $50 billion. Each is written so the fix is a control rather than an exhortation; the counts and dollars are Sable Creek’s where the audit found them.
| Failure | What it looks like | Why it matters | Fix |
|---|---|---|---|
| Inventory built from self-reporting | Business units list their models annually; a propensity model, a deposit-pricing optimizer and an OREO haircut tool absent for two years | Aggregate model risk is understated; unvalidated models drive decisions | Annual MRM discovery scan from the application inventory, contracts, GL and change tickets |
| Tiering by convenience | Four Tier 1 models downgraded to Tier 2 in 2025 as the validation backlog grew; no re-score on file | The largest exposures receive the least challenge | Published scoring criteria; downgrades need MRC approval with a re-score attached |
| Validation backlog without interim controls | 7 of 45 Tier 1 and 2 models overdue, one by nine months; decisions continue unchanged | The bank relies on models nobody has challenged within the policy window | Usage limits or overlays for overdue models; a dated recovery plan approved by the MRC; backlog reported to the board |
| Validation reports without numbers | Outcomes analysis reads “performance is acceptable”; no table, no threshold | Effective challenge has not occurred, whatever the cover says | Validation standard requires quantitative outcomes against thresholds set before testing |
| Monitoring without consequences | Threshold breached three consecutive quarters; report labeled informational; minutes record receipt only | The control exists on paper and the miss compounds | Monitoring standard defines the action on breach; breaches are a standing MRC item with recorded decisions |
| Findings closed by assertion; overlays sized by judgment | Validation condition closed on a management memo; $2.1 million overlay with no calculation while the backtest supported $4.4 million | The limitation persists behind a closed status; the allowance is wrong by a knowable amount | Closure requires validator re-test and MRC approval; overlays that answer a validation or monitoring result are anchored to it |
| Vendor model treated as validated by the vendor | “Vendor performs annual validation” recorded as the bank’s validation; no developmental evidence on file | The bank owns the risk and the examiner’s question | Bank-side conceptual review, customization testing and outcomes analysis; contractual right to vendor evidence |
| EUC purge after the 2026 definition | 31 spreadsheets removed from the inventory with no EUC register to receive them | Tools feeding pricing, valuations and reserves lose the only controls they had | EUC standard and register in place before reclassification |
| Model risk invisible to ERM | No operational loss event coded to model error; risk register rates model risk without indicators | Model failures never reach the board’s view of risk | Event type 7 coding for model losses; register entry with overdue validations, open highs and breaches as indicators |
The audit fails in its own recognizable ways. The first is re-validating: a team with one quant spends the budget reproducing the CECL platform’s cash flow engine, agrees with the vendor to the dollar, and reports nothing about the other 111 models, the inventory, or the closed finding that mattered. The second is auditing the binder: the policy is read, the org chart shows independence, four validation reports are “reviewed,” and no number in the file was computed by the auditor. The third is scoping from an untested inventory; the fourth is treating independence as a reporting-line question when the revised guidance has made it a rigor question; the fifth is accepting monitoring reports without treating them as IPE. The sixth is reporting a model miss as a model finding (“the PD is wrong”) instead of a control finding (“three controls saw the PD was wrong and none produced a decision”), which lets management fix the number and keep the process. Closure of every finding above should be validated on evidence, as the issue validation guide describes, and any that overlap an examination comment should track the MRA lifecycle so the bank reports one status for one problem.
Adapting the program: smaller banks, insurers, corporates and the committee
Sable Creek is a $9.6 billion bank, below the line above which the revised guidance is “most relevant,” and management raised that in the closing meeting. The answer is that the guidance was never the criterion; the policy is. The board adopted the framework in 2019 and reaffirmed it in June 2026 knowing the threshold, because a bank whose allowance, rate position and consumer credit decisions come out of models has significant model exposure whatever its asset size, and examiners will still assess those models under safety and soundness. What changes below $30 billion is proportion: fewer Tier 1 models, internal validation for Tier 2 with external validation reserved for the allowance and ALM models, a monitoring cadence matched to the reporting cycle, and a report that does not cite paragraph numbers from a document the agencies have said is not a rule. A bank with no MRM framework at all gets an advisory engagement first, to build the inventory, and the assurance engagement second.
Insurers run the same program with actuarial criteria. Actuarial Standard of Practice No. 56, Modeling, effective 1 October 2019, sets the actuary’s obligations on model understanding, validation, documentation and reliance on others, and it reads closely enough to the banking guidance that the work program transfers with vocabulary changes: pricing, reserving, capital and asset-liability models replace the credit and treasury models, and the appointed actuary’s opinion is read with the validation checklist. Corporates have no regulator asking, so the useful framing is financial reporting and decision reliance: the demand forecast that sets purchases, the rebate accrual model, the goodwill impairment cash flow and the pricing engine are models by any definition, and those feeding an accounting estimate already sit inside the ICFR program, where the estimate’s assumptions and the management review control over them are tested. The corporate inventory is the list of estimates and decisions that would move by more than a set dollar amount if the tool were wrong.
The audit committee wants three things: whether the models it relies on are being challenged, what the worst example means in dollars, and whether the fixes are controls or promises. Sable Creek’s engagement was rated Needs Improvement, with one High and three Medium findings dated inside 90 days; the CAE’s summary fit in a paragraph.
Audit committee summary. Model risk management at Sable Creek has the structure the policy describes: a complete inventory, re-performable tiering, and external validation of the eleven Tier 1 models. It does not yet have the reflexes. The CECL model’s commercial real estate default component has under-predicted defaults by 47 percent; the bank’s own monitoring flagged it in three consecutive quarters and its validator flagged it in 2024, and neither produced a decision. The allowance effect is $2.3 million to $4.4 million, and recalibration is agreed for the third-quarter close. The finding to track is the control one: monitoring breaches and validation findings must end in recorded decisions, and management has agreed the standard changes by 31 August.
Follow-up is where the engagement earns its rating. Standard 15.2 requires confirming that action plans were implemented: re-run the grade-level backtest on the recalibrated model after the September 2026 close, read the amended monitoring standard, and pull the MRC minutes for the first breach under it. The follow-up tests whether the framework now does what internal audit did.
Related guides
- Auditing AI and algorithms — fairness, explainability and generative-system controls.
- Operational risk: a comprehensive guide — Basel event types and where model failures are recorded.
- Internal audit in financial services: AML/KYC and compliance — transaction-monitoring scenarios and tuning.
- IRR vs IRRBB vs SIRR — the measures an ALM model produces and how to backtest them.
- Value at risk: the definitive guide — VaR methods and traffic-light backtesting.
- IPE testing — testing the extracts and reports monitoring depends on.
- Audit evidence — sufficiency and reliability when relying on validation work.
- Risk register — where aggregate model risk appears with indicators.
- Risk Library — the risk-by-risk guides hub.
- Topics — every guide by subject.
- Vendor and third-party models: validating what you cannot see inside
Leave a Reply