Most organizations now run decisions through models nobody on the audit committee has seen: a vendor’s résumé-ranking module that screens applicants before a recruiter looks, a credit-limit score that sets trade terms, an anomaly model that decides which employees get investigated, and a generative assistant into which staff paste whatever is on their screen. Each is a control, a risk, and a records problem at once, and each already carries live obligations: the EU AI Act’s prohibitions have applied since 2 February 2025 and its transparency duties since 2 August 2026, and New York City, Illinois, and Texas already regulate AI in hiring. The failure pattern is the same everywhere: management asserts the model is “monitored” and “human in the loop,” nobody can produce a complete inventory, and the only fairness evidence is a vendor’s attestation. An AI audit framework replaces those assertions with tests.
This guide gives you the working kit: a 24-row lifecycle framework of risks, controls, tests, and evidence; a fairness test worked through with confusion matrices by group, a computed disparate-impact ratio, and an equalized-odds gap; an AI inventory template with a risk-tiering rule set; a mapping of EU AI Act tiers to obligations and audit implications; and a generative-AI use-policy checklist. It was rewritten in September 2026 to reflect the Global Internal Audit Standards, NIST AI RMF 1.0, ISO/IEC 42001:2023, and the EU AI Act as amended in July 2026. It sits beside the model risk and algorithm audit guide, which covers SR 11-7 validation in depth, and the cybersecurity program audit guide, which covers the security controls this framework only references.
In this guide
- What an AI audit tests, and where the criteria come from
- Start with the inventory: template and risk-tiering rules
- The lifecycle audit framework: 24 rows of risks, controls, tests, and evidence
- Fairness testing with numbers: four-fifths screen, confusion matrices, equalized odds
- Explainability, human oversight, and the override test
- Regulatory mapping: EU AI Act tiers and US rules
- Third-party models and generative AI: vendor tests and a use-policy checklist
- Common failures in AI audits
- Scoping your first AI audit: budget, model language, and adaptations
What an AI audit tests, and where the criteria come from
An AI audit answers five questions with evidence rather than interviews. Who is accountable for the system, and did they approve it knowing what it does? Is the data it learned from, and the data it now sees, fit for the decision? Does it behave as claimed when accuracy, stability, and fairness are measured by someone other than the builder? Would anyone notice if it drifted, failed, or was attacked? Is the human oversight real, meaning a competent person with authority and time actually changes outcomes? Understanding the black box is not on the list: you do not need to read a model’s weights to test whether its selection rates differ by sex or whether the depot manager closes every flag in fourteen seconds.
Scope comes first, because “AI” covers four things that need different tests: deterministic rules and scoring formulas, tested by reperformance; models trained on history, tested by independent validation and outcome analysis; generative applications on a vendor’s general-purpose model, tested by output sampling, disclosure checks, and leakage tests; and agents that act across systems, tested through permission scopes and action logs. Under the Global Internal Audit Standards, Standard 13.4 requires the evaluation criteria to be named before fieldwork (the AI policy, the applicable law, NIST AI RMF, ISO/IEC 42001, and the model risk policy if one exists; where there is no AI policy, audit against the external frameworks and report the absence), and Standard 3.1 requires the competency to do the work: roughly 40 hours of a data scientist, co-sourced or borrowed under a written independence agreement from a team that did not build the systems under review.
NIST AI RMF, ISO/IEC 42001, the EU AI Act, and the IIA framework
Four documents supply the criteria. NIST AI RMF 1.0 (January 2023) is voluntary, with four functions, Govern, Map, Measure, and Manage, in 19 categories and 72 subcategories that read like control statements, which makes it the best source of test objectives; its Generative AI Profile (NIST AI 600-1, July 2024) adds twelve generative-AI risks, from confabulation to information security. ISO/IEC 42001:2023 is a certifiable management-system standard: clauses 4 through 10 mirror ISO/IEC 27001, and Annex A lists 38 controls under nine objectives, from AI policies (A.2) through impact assessment (A.5), the life cycle (A.6), data (A.7), and third parties (A.10); ISO/IEC 42006:2025 governs the certifying bodies and ISO/IEC 42005:2025 supplies an impact-assessment method. Where the organization is certified, the Statement of Applicability is your control catalog.
The EU AI Act (Regulation (EU) 2024/1689) is law rather than guidance and defines obligations by risk tier; the regulatory section gives the dates as amended from 27 July 2026. The IIA’s AI Auditing Framework, updated in 2024, is organized around the Three Lines Model with a practitioner’s guide of checklists; it is useful for structure and for language the audit committee already understands, but it contains no test procedures with numbers, which is why the lifecycle table exists. For banks and insurers, SR 11-7 (2011) remains the validation benchmark, and its conceptual soundness, ongoing monitoring, and outcomes analysis apply unchanged to machine-learning models.
| Audit domain | NIST AI RMF 1.0 | ISO/IEC 42001:2023 | EU AI Act (high-risk provisions) | Primary test in this guide |
|---|---|---|---|---|
| Accountability and policy | GOVERN 1.1, 2.1 | Clause 5; A.2.2, A.3.2 | Art. 17 (provider); Art. 26(1)–(2) (deployer) | Policy and charter inspection; trace two decisions through the committee |
| Inventory and classification | GOVERN 1.6 | Clause 4.1; A.4.2 | Art. 6 and Annex III; Art. 49 | Inventory reconciliation; re-tiering a sample |
| Data governance | MAP 2.3; MEASURE 2.10 | A.7.2–A.7.6 | Art. 10 | Lineage trace; representativeness comparison; label re-check |
| Fairness and bias | MEASURE 2.11 | A.5.4 | Art. 10(2)(f)–(g), 10(5) | Four-fifths screen with z-test; TPR and FPR by group |
| Transparency and explainability | MEASURE 2.8, 2.9 | A.8.2 | Art. 13; Art. 50; Art. 86 | Reason-code reproduction for 25 decisions |
| Human oversight | GOVERN 3.2; MANAGE 2.4 | A.9.2–A.9.4 | Art. 14; Art. 26(2) | Agreement-rate, time-per-case, and override analytics |
| Monitoring, logging, incidents | MEASURE 3.1; MANAGE 4.1, 4.3 | A.6.2.6, A.6.2.8, A.8.4 | Art. 12; Art. 26(5)–(6); Art. 72–73 | Drift-breach trace; reconstruction of 25 decisions from logs |
| Third parties | GOVERN 6.1; MAP 4.1 | A.10.2–A.10.4 | Art. 25; Art. 53 | Contract clause test; reperformance of vendor claims |
Start with the inventory: template and risk-tiering rules
The inventory is the population for everything else, and it is the artifact most likely to be wrong. MidState Beverage, the running example on this site (a three-state distributor with 12 depots, 300 delivery routes, and a six-person audit function), listed two systems when asked for its AI inventory during FY27 planning: a settlement-anomaly model finance had piloted at four depots after the FY27-01 route cash report, and a licensed generative assistant for 40 users. Discovery found nine; the other seven were the HR vendor’s applicant-ranking module inside the applicant tracking system, a routing optimizer in the fleet telematics platform, the ERP vendor’s cash-application matching model delivered in a 2025 upgrade, a customer-service chatbot on the ordering portal, and three consumer chatbots visible in DNS logs. An inventory completeness rate of 22 percent (2 of 9) is a finding before a single model is tested.
Discovery is a reconciliation against six sources: the single sign-on application list; procurement and purchasing-card spend for known AI vendors; cloud API keys and model-registry entries; a dependency scan of code repositories for machine-learning and LLM SDK imports; 90 days of DNS or cloud access security broker logs for generative-AI domains; and two years of ERP, HR, CRM, and telematics release notes, because vendors now ship models inside routine upgrades and nobody files a change request for them. In the template below, the later fields (population affected, degree of automation, human review point, jurisdiction) determine the tier, and the tier determines every downstream obligation. Register inventory completeness in the risk register as a control with a named owner.
Identity. System ID; name; business owner (accountable executive); technical owner; vendor and contract reference; type (rules, statistical or ML model, generative application, agent); version and deployment date; status (pilot, production, retired).
Purpose and use. Decision or task supported; population affected (customers, employees, applicants, the public, internal only); degree of automation (decides, recommends, assists); human review point and reviewer role; monthly volume; jurisdictions of use, including whether persons in the EU are affected.
Data. Input sources; personal data and special categories; training data sources and dates; lawful basis and data-sharing agreements; retention of inputs, outputs, and logs.
Risk and compliance. Risk tier (1–4) and the rule that assigned it; regulatory classification (EU AI Act tier and Annex III category, Article 6(3) exemption memo, state or sector rules); impact assessment date; known limitations; incidents in the last 12 months.
Assurance. Last validation date, scope, and validator; fairness metrics tested and results; explainability method; monitoring metrics, thresholds, and owner; retraining cadence and approval; last internal audit coverage; planned decommission date.
Tiering must be rule-based, because otherwise every team rates its own system “medium.” The trigger is the stakes of the decision and the people affected, never the sophistication of the technique.
| Rule | Trigger | Tier | Minimum assurance expected |
|---|---|---|---|
| 1 | Makes or materially influences a consequential decision about a person: employment, discipline, credit, insurance, housing, health, education, benefits | 1 | Independent validation before go-live and annually; fairness testing at each retrain; decision-level logging and explanations |
| 2 | Output feeds a key financial reporting control or an accounting estimate (cash application, allowance, forecasts used in judgments) | 1 if it feeds a key control or estimate; otherwise 2 | Design and operating effectiveness testing as for any key control; IPE testing on model outputs |
| 3 | Scores, ranks, or monitors employees’ behavior or performance | 1 | Worker notification; fairness by protected group; oversight analytics; in the EU, Annex III high-risk (emotion recognition prohibited except for medical or safety reasons) |
| 4 | Operational decisions at scale with no individual stakes (routing, demand forecasting, business-to-business pricing) | 2 | Validation at go-live; monitoring thresholds with an owner; annual performance review |
| 5 | Generative tool with access to non-public data | 2, rising to 1 if outputs reach customers, regulators, or filings without review | Approved-tool status; DLP; logging; sampled output review; enterprise terms excluding training on inputs |
| 6 | Assistive drafting with human review of every output | 3 | Policy acknowledgment, training, DLP, periodic output sample |
| 7 | Vendor-embedded model | Inherits the tier of the decision it supports, never the vendor’s label | Contract and documentation tests; reperformance of vendor claims on the organization’s own data |
| 8 | Any Tier 1 or 2 system used on persons in the EU | Unchanged, plus regulatory classification | Article 6 classification memo; Annex III category; Article 6(3) exemption documented and registered where claimed |
The lifecycle audit framework: 24 rows of risks, controls, tests, and evidence
The framework follows a system from the decision to build or buy it through to retirement. Each row names the risk in operational terms, the control a well-run organization would have, the test that produces evidence rather than a description, and the evidence to retain. A Tier 1 system gets every row; Tier 2 gets rows 1–5, 9–12, and 15–21; a Tier 3 tool gets rows 1, 3, 19, 20, and 22. Sample sizes follow the site’s 25/40/60 convention, and the rows load into the RCM Workbench as a risk-control matrix.
| # | Stage | Key risk | Control expected | Test performed | Evidence retained |
|---|---|---|---|---|---|
| 1 | Governance: accountability | No accountable executive; AI policy absent or aspirational | Board-approved AI policy; accountable executive per system; AI risk committee with charter | Inspect policy, charter, and 12 months of minutes; trace one approval and one incident | Policy, charter, minutes, RACI |
| 2 | Governance: regulatory scoping | Obligations unknown (EU AI Act tier, state statutes, sector rules) | Legal register of AI obligations mapped to inventory entries | Reperform classification for five systems; compare to the register | Classification memos, Article 6(3) documentation |
| 3 | Inventory | Inventory incomplete; shadow AI in business units | Mandatory registration before use; quarterly discovery routine | Reconcile inventory to the six discovery sources; compute completeness | Reconciliation workpaper, exceptions list |
| 4 | Intake and tiering | Systems tiered by builder, not by impact | Rule-based tiering with independent sign-off | Re-tier ten systems; challenge any Tier 2 or 3 system that decides about people | Tiering rules, intake forms, results |
| 5 | Use-case approval | Deployment without impact assessment | Impact assessment (ISO/IEC 42005 or Article 27 style) before build or purchase | For five systems compare assessment date to go-live date; read for population affected | Impact assessments, approvals |
| 6 | Data sourcing and lawful basis | Personal data without lawful basis; data of unknown provenance | Data protection assessment; lawful basis recorded; agreements for external data | Trace each training dataset to source, lawful basis, and contract; check special-category data | Assessment, lineage records, contracts |
| 7 | Data quality and representativeness | Training population differs from deployment population; label errors | Data profiling; representativeness analysis; label quality assurance | Compare training and production distributions on five variables; re-check 60 labels against source records | Datasheets, profiling reports, label workpaper |
| 8 | Feature engineering | Proxies for protected attributes (postal code, employment gaps, name-derived features) | Feature review for proxies with documented exclusions | Correlate each feature with protected attributes on the validation sample; obtain the exclusion log | Feature dictionary, correlation table |
| 9 | Development environment | Uncontrolled code; results not reproducible | Version control, experiment tracking, model registry | Re-run training for one model from the registry entry; compare to the registered version | Repository history, reproduction results |
| 10 | Independent validation | Validation by the builder; no challenger | Validation independent of development, scoped by tier, with findings log (SR 11-7 pattern) | Read the validation report; confirm reporting-line independence; reperform two tests | Validation report, findings log |
| 11 | Performance testing | Accuracy claimed on training data only | Out-of-sample and out-of-time tests against approved thresholds | Reperform metrics on the holdout; compare to thresholds and go-live approval | Holdout definition, results, approval |
| 12 | Fairness testing | No metric defined, or the wrong one for the decision | Metrics and thresholds per use case; trade-offs documented; retest at each retrain | Compute impact ratio with z-test and TPR/FPR by group (worked example below) | Fairness workpaper |
| 13 | Explainability | No usable explanation; reason codes unrelated to score drivers | Method appropriate to the model; reason codes validated against features | For 25 decisions reproduce the explanation; check reason codes against top contributing features | Explanation samples, reproduction results |
| 14 | Security and adversarial robustness | Prompt injection, data poisoning, model extraction, unprotected artifacts | Threat model; red-team testing; input validation; artifact access controls | Inspect the red-team report; run five injection cases against any chatbot; review artifact access | Threat model, red-team results, test log |
| 15 | Deployment and change management | Untested changes; silent retraining; no rollback | Change approval; staged rollout; rollback plan; retraining treated as a change | Trace 40 model version changes to approvals, tests, and release notes | Change tickets, release notes |
| 16 | Access management | Broad access to training data, models, prompts, pipelines | Role-based access; privileged access reviews; secrets management | Pull access listings for registry, feature store, prompt repository, and pipeline; test appropriateness (see the IAM audit guide) | Access listings, review evidence |
| 17 | Logging and record-keeping | Decisions cannot be reconstructed | Automatic logging of inputs, outputs, version, reviewer (Article 12); deployer retention of at least six months (Article 26(6)) | Select 25 decisions from business records; reconstruct each from the logs | Log extracts, reconstruction workpaper |
| 18 | Monitoring and drift | Drift undetected; thresholds absent; alerts ignored | Monitoring metrics with thresholds, owners, response times | Obtain six months of monitoring output; list breaches and responses; recompute one metric from raw data as IPE testing | Dashboards, alert tickets, recomputation |
| 19 | Human oversight | Rubber-stamping; automation bias; reviewers without authority | Defined review points; competence requirements; override rights; oversight metrics reported | Analyze review logs for agreement rate, time per case, override rate by reviewer, override outcomes | Review logs, oversight analytics |
| 20 | Incident management | AI incidents not recognized; regulatory reporting missed | AI incident definition; escalation path; serious-incident reporting where required (Article 73) | Inspect the incident log; test three incidents end to end; search complaints for AI-related causes | Incident log, root causes, complaint search |
| 21 | Third-party models and vendors | Vendor claims untested; contract gaps; unnoticed model changes | Due diligence; contract terms on data use, change notice, documentation, test rights | Inspect five AI vendor contracts; obtain evidence of change notifications; compare SOC 2 scope to model claims | Contracts, due-diligence files, clause matrix |
| 22 | Generative AI use | Data leakage; fabricated outputs used in decisions | Approved-tool list; DLP; usage policy; output verification for high-stakes uses | Run the use-policy checklist below; sample 25 outputs used in decisions | DLP logs, attestations, output sample |
| 23 | AI literacy and training | Overseers cannot operate or challenge the system (Article 4; Article 26(2)) | Role-based training; competence records for oversight roles | Test completion for oversight roles; interview five overseers on what they would do with an anomalous output | Training records, interview notes |
| 24 | Decommissioning | Retired models still scoring; artifacts lost before disputes are resolved | Decommission procedure; artifact archive aligned to retention and legal hold | Trace three retired systems to shutdown evidence and archived artifacts | Decommission records, archive index |
Two rows are where most first engagements fail. Row 17: pick 25 real decisions from the business side and ask for the model version, inputs, output, and reviewer for each; if they cannot be reconstructed, every other assertion about the system is untestable. Row 18: a monitoring dashboard is information produced by the entity, and its completeness and accuracy have to be tested before you rely on it, exactly as for a management review control in the close.
Fairness testing with numbers: four-fifths screen, confusion matrices, equalized odds
MidState’s driver-hiring funnel supplies the worked example. In FY26 the company received 3,000 applications for CDL delivery-driver openings. The applicant tracking system’s vendor-supplied ranking module scores each application from 0 to 100 and recommends an interview at 62 or above; the ATS event log shows 91 percent of interviews came from the recommended pile. The vendor’s bias statement says the tool was “tested for demographic parity on the training corpus,” which describes the vendor’s data, not MidState’s applicants. The test therefore runs on MidState’s own population in four steps: selection rates, statistical significance, error rates by group, and cause.
Step 1: selection rates and the four-fifths screen
| Group | Applications | Recommended for interview | Selection rate | Impact ratio |
|---|---|---|---|---|
| Men (reference group, highest rate) | 2,600 | 1,170 | 45.0% | 1.00 |
| Women | 400 | 136 | 34.0% | 0.756 |
| All applicants | 3,000 | 1,306 | 43.5% | n/a |
The impact ratio is the selection rate of the group under review divided by the highest group’s rate: 34.0 ÷ 45.0 = 0.756. The four-fifths rule in the Uniform Guidelines on Employee Selection Procedures (29 CFR 1607.4(D), 1978) treats a ratio below 0.80 as evidence of adverse impact, and the Guidelines remain in force even though the EEOC removed its AI-specific technical assistance from its website in January 2025. The screen is sensitive to small groups: nine more recommended women (145 rather than 136) would have lifted the ratio to 0.806, and the Guidelines themselves say small differences can be adverse impact if statistically significant while large ones may not be if they rest on small numbers. Step 2 is therefore not optional.
Step 2: is the difference real?
A two-proportion z-test settles the small-numbers objection. The pooled selection rate is 1,306 ÷ 3,000 = 0.435; the standard error of the difference is the square root of 0.435 × 0.565 × (1/2,600 + 1/400) = 0.0266; the observed difference of 0.110 divided by 0.0266 gives z = 4.13, a p-value below 0.0001. Courts and enforcement agencies commonly treat two to three standard deviations as significant; this is over four, so the screen and the test agree. Had the ratio been 0.78 with z = 1.2, the honest conclusion would have been “screen failed, difference not significant, retest next quarter on a larger population,” and the workpaper should say exactly that.
Step 3: confusion matrices by group
Selection rates show the outcome differs; they do not show whether the tool makes different errors. That needs a ground-truth label for both recommended and non-recommended applicants, which ordinary hiring never produces because rejected applicants are never evaluated. At the audit’s request, a two-recruiter panel reviewed a stratified sample of 800 FY26 applications blind to the tool’s score, using the structured qualification checklist for the role (valid CDL, clean motor vehicle record, verifiable driving experience, physical requirements); women were oversampled to 300 of the 800. The panel found similar qualification rates, 52 percent of sampled men and 50 percent of sampled women, so the selection gap cannot be explained by the applicant pools.
| Group (sample) | Qualified, recommended (TP) | Qualified, not recommended (FN) | Not qualified, recommended (FP) | Not qualified, not recommended (TN) | True positive rate | False positive rate | Precision (PPV) |
|---|---|---|---|---|---|---|---|
| Men (n = 500) | 208 | 52 | 36 | 204 | 80.0% | 15.0% | 85.2% |
| Women (n = 300) | 96 | 54 | 15 | 135 | 64.0% | 10.0% | 86.5% |
| Gap (men minus women) | n/a | n/a | n/a | n/a | 16.0 points | 5.0 points | −1.3 points |
Read the last three columns. The true positive rate is the share of qualified applicants the tool recommends: 80 percent of qualified men and 64 percent of qualified women, so a qualified woman is screened out 36 percent of the time against 20 percent for a qualified man. That is a failure of equal opportunity, and the 16-point gap is far beyond the 5-point tolerance most fairness policies set. Equalized odds requires both true positive and false positive rates to match, and the tool fails that too: it is stricter with unqualified women (10 percent let through) than with unqualified men (15 percent). Precision is almost identical at 85.2 and 86.5 percent, which is what the vendor means by “equally accurate across groups”: predictive parity can hold while equal opportunity fails. Choose the metric from the harm, not from what the vendor reports: for a screening tool it is the rate at which qualified people are wrongly excluded, for a fraud or monitoring model it is the false positive rate, and where scores are used as probabilities it is calibration by score band.
Step 4: cause, the less discriminatory alternative, and the finding
Cause analysis starts with the feature dictionary. Two of the tool’s 31 features carried most of the gap: “months of continuous employment in the last five years” and a “gap penalty” for breaks longer than six months. Women’s median employment gap in MidState’s pool was 14 months against 3 for men, so the features functioned as proxies for sex, and neither had been validated against job performance, which the Uniform Guidelines require of any procedure with adverse impact (29 CFR 1607.5). The vendor re-scored the FY26 population without the two features: women’s true positive rate rose to 74.7 percent against 79.2 for men (a 4.6-point gap), false positive rates were 11.3 and 14.6 percent, the population impact ratio became 0.883 (selection rates of 44.2 and 39.0 percent), and the holdout AUC fell from 0.81 to 0.80. The residual difference had z = 1.95, at the edge of significance, so the workpaper recorded the alternative as materially less discriminatory and set a retest for the next quarter. The organization cannot instead use a different cutoff for women, because Title VII section 703(l) prohibits adjusting scores or cutoffs on the basis of sex or race; in the EU, Article 10(5) of the AI Act permits processing special-category data for this kind of bias detection. The finding follows the 5 Cs pattern and is rated High under the severity rubric.
Condition. The applicant-ranking module recommended 34.0 percent of women applicants for interview against 45.0 percent of men in FY26 (impact ratio 0.756; z = 4.13). A blind panel review of 800 applications found qualified women were not recommended 36 percent of the time against 20 percent for qualified men. Two unvalidated features acted as proxies for sex.
Criteria. Title VII and 29 CFR 1607.4(D) and 1607.5; the AI policy requirement that Tier 1 systems be fairness-tested on company data before use; NIST AI RMF MEASURE 2.11.
Cause. The tool was procured as an ATS feature and never entered the AI inventory; HR relied on the vendor’s training-corpus bias statement; no owner was assigned to test outcomes on MidState applicants.
Effect. Potential disparate impact on a protected basis across roughly 400 applications a year, with legal and reputational exposure, and an estimated 50 qualified applicants a year excluded from a role the company struggles to fill.
Recommendation. Suspend the two proxy features immediately; complete a validation study of the remaining features within 120 days; add the tool to the inventory as Tier 1 with quarterly impact-ratio and TPR-by-group reporting to the AI risk committee; amend the vendor contract to require subgroup performance reporting and change notification.
Explainability, human oversight, and the override test
Explainability has three audiences. The affected person needs reasons: Regulation B requires specific principal reasons in credit adverse-action notices (12 CFR 1002.9), Article 86 of the EU AI Act gives people subject to certain high-risk decisions a right to a clear explanation of the system’s role and the main elements of the decision, and Colorado’s SB 26-189, in force from 1 January 2027, requires a plain-language explanation within 30 days of an adverse outcome. The overseer needs to know what to look at before agreeing or overriding, and the validator needs global behavior. Post-hoc tools such as SHAP and LIME serve the validator and, with care, the overseer; they are local approximations, and they explain the model rather than the decision policy, because the threshold and the override sit outside the model. The test is reproduction: take 25 decisions and check that each explanation can be regenerated from the logged inputs and model version, that the stated reasons match the features that actually moved the score, and that catch-all reasons such as “other” appear in fewer than 5 percent of cases. An explanation you cannot regenerate is a representation, not audit evidence.
Human oversight is a management review control and is tested the same way, by measuring whether the reviewer could have caught an error and whether they ever do. MidState’s settlement-anomaly model is the cautionary case. Finance piloted it at four depots (100 routes) for 24 weeks, scoring each driver settlement from 0 to 100 and flagging scores of 70 or above for depot manager review. The pilot covered 12,600 settlements and produced 478 flags, a 3.8 percent flag rate. Managers closed 471 of the 478 as “reviewed, no exception,” a 98.5 percent agreement rate; the median time from opening to closing a flag was 14 seconds, and 61 percent of closures fell in a single 15-minute window on Monday mornings. Seven flags became confirmed shortages totaling $2,310, so the model was finding something. Two managers closed every flag on routes for which they also approve variance overrides, the self-approval weakness the FY27-01 report had already reported. The threshold of 70 had been chosen because it kept flags below 4 percent of settlements, “what the depots said they could review”: the model was tuned to reviewer capacity rather than to the cost of a missed shortage.
Four metrics make oversight auditable and belong in management’s own reporting: agreement rate by reviewer (above 95 percent with no documented reasons is a red flag, not a compliment), the distribution of time per case, override rate by reviewer with the recorded reason, and the outcome of overrides. Add competence and authority, which Articles 14 and 26(2) of the EU AI Act require and which any policy that calls the reviewer a control implies. For the settlement model the fix was a threshold set from a cost-of-miss analysis, independent review of flags where the manager approves overrides, mandatory closure reasons, and a monthly report with the four metrics, which is now also how the fraud red flags from the route cash work are tracked.
Regulatory mapping: EU AI Act tiers and US rules
As of September 2026 the EU AI Act’s timetable is the amended one. The Digital Omnibus on AI, Regulation (EU) 2026/1744, was adopted on 8 July 2026, published on 24 July, and entered into force on 27 July 2026. It moved the application date for stand-alone high-risk systems in Annex III from 2 August 2026 to 2 December 2027, and for high-risk AI embedded in Annex I products from 2 August 2027 to 2 August 2028. It left the prohibitions (since 2 February 2025), the general-purpose AI model obligations (since 2 August 2025), and the Article 50 transparency duties (since 2 August 2026) on their original dates, with a grace period to 2 December 2026 for machine-readable marking by systems already on the market; it softened Article 4 from ensuring AI literacy to supporting it and extended SME relief to small mid-cap companies. Penalties are unchanged: up to €35 million or 7 percent of worldwide turnover for prohibited practices, €15 million or 3 percent for most other breaches, and €7.5 million or 1 percent for supplying incorrect information. For an audit function that means fifteen months of runway for high-risk systems and none for chatbot disclosure and content marking.
| Tier | Examples an auditor will meet | Key obligations | Applies from (as of September 2026) | Audit implication |
|---|---|---|---|---|
| Prohibited (Article 5) | Emotion recognition of employees or students except for medical or safety reasons; social scoring; manipulative techniques; untargeted facial-image scraping | Cannot be placed on the market or used | 2 February 2025; new imagery prohibition transitional to 2 December 2026 | Screen HR, fleet, and customer-experience technology; driver-facing fatigue cameras need a documented safety rationale |
| High-risk, Annex III (stand-alone) | Recruitment screening and ranking; promotion, termination, and task-allocation tools; worker monitoring; credit scoring of natural persons (except fraud detection); life and health insurance pricing; education; essential public services; biometrics; critical infrastructure; law enforcement; justice | Providers: Art. 9–15 (risk management, data governance, documentation, logging, instructions, human oversight, accuracy and security), Art. 17, 43, 49, 72–73. Deployers: Art. 26 (competent oversight, six-month log retention, informing workers’ representatives and affected persons), Art. 27 fundamental rights impact assessment for public bodies and credit and insurance deployers, Art. 86 explanations | 2 December 2027 (moved from 2 August 2026) | Classification memo per system including any Article 6(3) exemption; gap assessment against Art. 26 (deployers) or Art. 9–17 (in-house builds); readiness plan reported to the committee now |
| High-risk, Annex I (product-embedded) | AI safety components in machinery, medical devices, vehicles, and other regulated products | Conformity through the sector product legislation plus AI Act requirements | 2 August 2028 (moved from 2 August 2027) | Confirm product compliance teams own the AI requirements |
| Transparency (Article 50) | Customer chatbots; AI-generated marketing images and text; deepfakes; emotion recognition where permitted | Tell people they are interacting with AI; mark synthetic content in machine-readable form; disclose deepfakes | 2 August 2026; marking grace to 2 December 2026 for systems on the market before 2 August 2026 | Test the chatbot for disclosure at first interaction; inspect marking configuration; sample published marketing assets |
| General-purpose AI models (Articles 53–55) | The foundation models behind licensed assistants and in-house generative applications | Provider obligations: technical documentation, downstream information, copyright policy, training-content summary; extra duties for systemic-risk models | 2 August 2025; models on the market before that date by 2 August 2027 | The deployer’s test is whether the vendor documentation exists and feeds the organization’s impact assessments |
| Minimal risk (everything else) | Routing optimizers, demand forecasts, spam filters, internal drafting tools | No specific obligations beyond Article 4 literacy measures; voluntary codes | Article 4 since 2 February 2025 | Tier under the internal rule set anyway; the Act is a floor |
The United States has no federal AI statute as of September 2026. Executive Order 14365 of 11 December 2025 directed the Department of Justice to form a task force to challenge state AI laws, but it does not itself preempt anything, and the state laws below remain on the books; verify each before you cite it. New York City’s Local Law 144 has required an annual independent bias audit of automated employment decision tools since 5 July 2023, with impact ratios by sex and race or ethnicity published on the employer’s site. Illinois HB 3773 amended the Illinois Human Rights Act from 1 January 2026 to cover discriminatory use of AI in employment, Texas’s TRAIGA took effect the same day with an intent standard, and California’s civil rights regulations on automated decision systems in employment have applied since 1 October 2025. Colorado repealed its 2024 AI Act before it took effect and replaced it with SB 26-189, signed on 14 May 2026 and effective 1 January 2027, which drops the duty of care and impact assessments in favor of pre-use notice, adverse-outcome explanation within 30 days, a human review path, and three-year records. Sector rules did not wait: Regulation B adverse-action reasons, the Fair Credit Reporting Act, SR 11-7, and the GDPR’s Article 22 limits on solely automated decisions all apply today, and the test under each is whether the organization can show the notice, the reasons, the human intervention path, and the records for 25 real decisions.
Third-party models and generative AI: vendor tests and a use-policy checklist
Most AI in a non-technology company arrives from vendors, and the assurance most vendors offer is a SOC 2 report, which covers the service’s security, availability, confidentiality, processing integrity, and privacy and says nothing about whether the ranking module screens out qualified women. The documentation to demand is the model card or equivalent (intended use, training data categories, known limitations, performance by subgroup), a change-notification commitment, data-use terms (no training on your inputs, retention, sub-processors), incident notification, test rights, and exit terms that return logs and artifacts. The IIA’s Third-Party Topical Requirement, effective 15 September 2026, makes governance, risk management, and control requirements mandatory whenever third-party management is in scope, and vendor AI is the cleanest case for it; under the EU AI Act a deployer that substantially modifies a high-risk system becomes its provider (Article 25), and Colorado’s SB 26-189 obliges developers to supply deployers with documentation.
Five tests cover vendor models. Inspect five AI vendor contracts against a clause matrix (data use, change notice, documentation, subgroup performance, test rights, incident notice, exit) and count the gaps. Obtain evidence that change notifications were received and acted on in the last year; silence usually means the model changed and nobody noticed. Reperform the vendor’s headline claims on the organization’s own data, as the fairness section did. Map the data flows: what leaves, where it is processed, how long it is kept, and which sub-processors see it. Finally, assess exit: if the vendor withdrew the model tomorrow, could the process run, and would decision logs survive for a dispute two years from now? Cloud-hosted models add the shared-responsibility questions in the cloud audit guide.
Generative AI use is a policy-compliance audit with a data-leakage core, and MidState’s discovery data shows the shape of it. The 90-day DNS and CASB extract showed 214 uploads to consumer chatbot domains by 31 users, three of which matched the naming pattern of the customer master files for the two acquired distributors, files holding roughly 1,900 customer records including ACH bank details. The acceptable-use policy had existed since March 2026, but acknowledgment was not tracked, DLP had not been configured for generative-AI domains, and the consumer tools ran on terms that let the provider train on inputs unless the user opts out; the licensed enterprise assistant excluded that, which is the reason to pay for it. Each checklist row has a pass criterion so the result is a count, not an impression.
| # | Policy requirement | Test | Evidence | Pass criterion |
|---|---|---|---|---|
| 1 | Approved-tool list with tiers of permitted data | Compare the list to 90 days of DNS and CASB traffic to generative-AI domains | Traffic extract, approved list | At least 95 percent of sessions on approved tools; other domains blocked or under exception |
| 2 | Data classification rules for prompts and uploads | Seed a file with a confidential marker; attempt upload to an approved and an unapproved tool | DLP configuration, test log | Upload blocked or logged with an alert within the response time |
| 3 | Enterprise terms for approved tools | Inspect the contract or terms for each approved tool | Contracts, terms of service | Explicit exclusion of training on inputs; configurable retention; sub-processors listed |
| 4 | Identity, licensing, and leavers | Reconcile licensed users to the identity provider; test 40 leavers | License and IAM extracts | No active licenses for leavers; SSO and MFA enforced (see the user access review guide) |
| 5 | Output verification before use in decisions or external documents | Sample 25 outputs used in finance, legal, or customer communications; obtain review evidence | Documents, review evidence | 100 percent reviewed for Tier 1 uses, reviewer named |
| 6 | Incident reporting for data entered into tools | Compare the incident log with DLP and DNS evidence of uploads | Incident log, DLP alerts | Every DLP alert has an incident record and a closed response |
| 7 | Customer-facing disclosure | Open the customer chatbot; check for disclosure at first interaction | Screenshots, configuration | Disclosure before the first substantive response (Article 50(1)) |
| 8 | Synthetic-content marking | Inspect marking configuration for generators used in marketing; sample ten published assets | Configuration, asset sample | Machine-readable marking in place (Article 50(2); grace to 2 December 2026 for pre-existing systems) |
| 9 | Training and acknowledgment | Test completion for all users of approved tools and everyone in an oversight role | LMS records, acknowledgments | At least 95 percent within 30 days of access |
| 10 | Agent permissions and prompt-injection resistance | Review permission scopes granted to agents; run five injection cases against agents that can write to systems | Scope listings, test results | No unbounded write scopes; injection cases fail to execute actions |
Common failures in AI audits
These are the failures I see in first and second AI engagements, on both sides of the table. Most come from importing the shape of a policy-compliance review into a subject that needs outcome testing.
| Failure | What it looks like | Why it matters | Fix |
|---|---|---|---|
| Auditing the policy, not the systems | A report confirming a policy, a committee, and training exist, with no system named | The committee can read the policy itself; assurance means someone tested a model | Sample three systems by tier and run the lifecycle rows; report per system |
| Accepting vendor bias attestations | “The vendor tested for bias” filed as evidence with no numbers and no population | Vendor tests run on vendor data; adverse impact is a property of your applicants and your threshold | Reperform selection and error rates on the organization’s own population |
| Taking the inventory as given | Fieldwork scoped to the two systems management listed | The untested systems are the risk; vendor-embedded ones are almost never listed | Run the six-source discovery reconciliation before scoping |
| Wrong metric, no significance test | An impact ratio of 0.83 reported as “passes,” or 0.76 as “fails,” with 80 people in the smaller group | Small groups swing the ratio; the wrong metric hides the harm that matters | Pair the four-fifths screen with a z-test; choose TPR or FPR gaps from the harm |
| “Human in the loop” accepted as a control | A review step on the process map; nobody measured agreement rates or time per case | A 98.5 percent agreement rate in 14 seconds is rubber-stamping, and it is the norm | Compute the four oversight metrics from the review log |
| Monitoring reports relied on without IPE testing | Drift dashboard screenshots filed as evidence | Reports the model team builds from its own pipeline are information produced by the entity | Recompute one metric from raw data; test report completeness and parameters |
| Findings closed without retest after retraining | Remediation evidence is a new model version; nobody re-ran the fairness test | Retraining changes behavior; the original finding may have returned | Treat retraining as a change; require retest evidence in issue validation |
| Auditors’ own analytics held to a lower standard | A team-written scoring script selects samples with no validation or version control | The function loses standing to ask for controls it does not apply itself | Apply rows 9, 11, and 15 to your own tools, as the journal entry analytics guide does for scoring scripts |
Scoping your first AI audit: budget, model language, and adaptations
First engagements take one of two shapes. A program audit covers governance, the inventory reconciliation, and a sample of three systems across tiers, and produces the map the committee needs to fund the rest. A single-system deep dive takes one Tier 1 system through every lifecycle row with a full fairness test, and is the right choice when a system is already drawing regulatory or legal attention. The budget assumes a mid-sized organization, an audit team without a data scientist, and a co-sourced specialist; it is generous on inventory and reporting because those are where first engagements overrun. Build it in the work program format and write the criteria into the planning memo before fieldwork.
| Phase | Program audit (hours) | Single Tier 1 system (hours) | Who |
|---|---|---|---|
| Planning, criteria, and stakeholder interviews | 40 | 30 | Audit manager, senior |
| Inventory discovery and reconciliation | 60 | 10 | Senior, IT audit |
| Governance, policy, and regulatory scoping | 40 | 15 | Audit manager |
| System testing (lifecycle rows) | 135 (three systems) | 120 | Senior, IT audit |
| Fairness and outcome testing | 40 | 60 | Senior plus co-sourced data scientist |
| Generative AI and third-party tests | 40 | 20 | IT audit |
| Reporting, committee materials, quality review | 45 | 35 | Audit manager, CAE |
| Co-sourced specialist (validation reperformance, feature analysis) | 40 | 40 | External or independent internal data scientist |
| Total | 440 | 330 | n/a |
Objective and scope language matters more than usual, because auditees will try to narrow the engagement to “governance” and the committee will assume it covers every model in the company. The planning-memo language below is for a program audit.
Objective. To assess whether governance over AI and algorithmic systems is designed and operating to identify the systems in use, classify them by risk and regulatory obligation, and ensure that systems making or influencing consequential decisions are validated, monitored, fairness-tested, and subject to effective human oversight.
Scope. The AI governance framework and inventory as at 30 September 2026; inventory completeness tested against six discovery sources; detailed testing of three systems selected by tier: the applicant-ranking module (Tier 1, employment), the settlement-anomaly model (Tier 1, employee monitoring), and the licensed generative assistant (Tier 2). Period: 1 October 2025 to 30 September 2026.
Criteria. The AI Acceptable Use Policy (March 2026); NIST AI RMF 1.0; ISO/IEC 42001:2023 Annex A; Title VII and 29 CFR Part 1607 for the employment system; the EU AI Act where persons in the EU are affected; the model risk policy for validation standards.
Out of scope. Technical review of model architecture; the vendor’s internal controls beyond contract and documentation obligations; the ERP cash-application model, covered in the FY27 ICFR support engagement.
Adapt the framework to the organization rather than the other way round. A small company with no data scientists has almost entirely vendor AI, and the right first engagement is the inventory reconciliation, the generative-AI checklist, and one vendor-model deep dive, at about 120 hours; the guide to establishing a function in a small company covers staffing it. A bank or insurer already has model risk management under SR 11-7, so the engagement audits that function’s coverage: does the inventory include the ML and generative systems outside the model shop, and do validations test fairness and drift or only conceptual soundness? EU deployers of credit and insurance systems and public bodies add the fundamental rights impact assessment (Article 27), registration (Article 49), and worker notification (Article 26(7)). An agent with a write scope is a privileged user, and the privileged access audit procedures apply to it unchanged.
When you present results, lead with people and money rather than metrics: “one in three qualified women applicants was screened out by a tool nobody had tested, and the fix costs nothing” lands, and “equalized odds was violated” does not. Report inventory completeness as a percentage, oversight as an agreement rate and a median review time, and regulatory readiness as a count of high-risk systems with and without a classification memo, and put retest dates in the issue log so that retraining cannot quietly close a finding. AI coverage belongs in the audit plan as a recurring line. The Topics hub and Risk Library hold the related risk entries, the Tools page has the free sampler and RCM builder used here, and the By Role page maps this material by role.
Related guides
- Model risk and algorithm audit — SR 11-7 validation in depth.
- IPE testing — testing monitoring reports and model outputs before relying on them.
- Third-Party Topical Requirement — mandatory from 15 September 2026 for vendor AI in scope.
- Auditing cybersecurity programs — the controls behind the security row.
- IAM audit — access testing for registries, feature stores, prompt repositories, and agents.
- Journal entry analytics — controls for the audit function’s own scoring scripts.
- Global Internal Audit Standards — Standards 3.1 and 13.4 on competency and criteria.
- IIA Topical Requirements — the requirements that intersect with AI engagements.
- Continuous auditing vs continuous monitoring — where model monitoring belongs.
- All guides — the complete catalog.
Leave a Reply