Most business continuity plans are tested the day they are needed, and most fail the test in ways the plan never contemplated: the backup restores but takes 40 hours instead of four, the alternate site has desks but no access to the ERP, the call tree reaches people who left two years ago, the critical supplier’s own plan turns out to be “call our account manager.” Internal audit has historically checked that a plan exists, that it was updated this year, and that a tabletop exercise was held, and has signed off. That is an audit of documents. The organizations that survived the last decade’s ransomware attacks, cloud outages, and supplier collapses were not the ones with the best documents; they were the ones that knew which processes could not stop, for how long, and had proven, not asserted, that they could bring them back.
This guide sets out how to audit business continuity as a capability rather than a binder, and it is written to the IIA’s Organizational Resilience Topical Requirement, issued 30 April 2026 and mandatory from 30 April 2027 for any engagement in which resilience is in scope. It covers the 2026 regulatory setting (the Topical Requirement, DORA, the UK operational resilience regime, ISO 22301), the vocabulary that has to be precise (BIA, RTO, RPO, MTPD, impact tolerance), all twenty requirements of the Topical Requirement with the evidence for each, how to audit a business impact analysis, a criticality tier model, a testing hierarchy and what a credible test program looks like, the third-party and technology dependencies where most plans break, a 380-hour work program, a worked engagement at MidState Beverage whose route-accounting module had a 72-hour recovery target for a process that cannot stop for eight, a metrics set for the board, and the mistakes that turn resilience audits into document reviews. It was rewritten in September 2026. The site’s guides to auditing cybersecurity programs and cloud risk cover the two dependencies this guide treats as inputs.
In this guide
- The 2026 setting: the Topical Requirement, DORA, UK operational resilience, ISO 22301
- The vocabulary that has to be precise
- The twenty requirements and the evidence for each
- Auditing the business impact analysis and the criticality tiers
- Testing: the hierarchy, a credible program, and how tests fail
- Where plans break: third parties, technology, backups, and people
- The work program
- Worked example: MidState Beverage’s route accounting module
- Metrics for the board
- Common mistakes in continuity audits
The 2026 setting: the Topical Requirement, DORA, UK operational resilience, ISO 22301
Four sources define what a resilience audit is measured against in 2026, and an auditor should know which apply to the organization before scoping. The IIA’s Organizational Resilience Topical Requirement, issued 30 April 2026 and effective 30 April 2027, is the fourth in the series after Cybersecurity, Third-Party Risk Management, and Organizational Behavior; it defines organizational resilience as the ability of an organization to absorb and adapt in a changing environment, scopes it across business continuity, disaster recovery, crisis management, and operational resilience in their operational, technological, and financial dimensions, and sets out requirements the auditor must assess, or document a rationale for excluding, whenever resilience is in scope. For EU financial entities and their critical ICT providers, the Digital Operational Resilience Act has applied since 17 January 2025, with its requirements for ICT risk management, incident reporting, resilience testing including threat-led penetration testing for the largest firms, and a register of ICT third-party arrangements. In the United Kingdom, the PRA and FCA operational resilience regime required firms to identify important business services, set impact tolerances, and be able to remain within them by the end of the transition period on 31 March 2025, which is the regime that introduced the impact-tolerance concept the rest of the profession is now borrowing. And ISO 22301:2019 remains the management-system standard for business continuity that most non-financial organizations certify to or benchmark against, with its cycle of policy, business impact analysis, risk assessment, strategy, plans, exercising, and review. US banks additionally have the 2020 interagency Sound Practices to Strengthen Operational Resilience, which applies formally to the largest institutions and informally to everyone their examiners visit.
| Source | Applies to | Core concept | What it changes for the audit |
|---|---|---|---|
| IIA Organizational Resilience Topical Requirement (issued 30 Apr 2026, effective 30 Apr 2027) | Every internal audit function conforming to the Global Internal Audit Standards, whenever resilience is in scope | Twenty requirements across governance, risk management, and controls; documented applicability assessment | The criteria set is fixed; exclusions must be justified in the file |
| DORA (EU Regulation 2022/2554, applying from 17 Jan 2025) | EU banks, insurers, investment firms, payment institutions, crypto-asset providers, and their critical ICT third-party providers | ICT risk management framework; major incident reporting; digital operational resilience testing; ICT third-party risk and the register of information | Resilience of ICT becomes a supervised obligation with specific artifacts (the register, test reports, incident classifications) to audit |
| UK operational resilience (PRA SS1/21, FCA PS21/3; transition ended 31 Mar 2025) | UK banks, building societies, insurers, investment firms, payment and e-money firms | Important business services; impact tolerances; mapping and scenario testing; self-assessment | The audit tests whether the firm can stay within tolerance, not whether it has a plan |
| ISO 22301:2019 | Any organization, by choice or contract | Business continuity management system: BIA, risk assessment, strategies, plans, exercises, continual improvement | A recognized benchmark for structure and documentation; certification audits are not assurance of capability |
| US interagency Sound Practices (Oct 2020) and NIST SP 800-34 (IT contingency planning) | Largest US banks formally; widely used as benchmarks | Resilience of critical operations; IT contingency planning lifecycle | Vocabulary and expectations examiners bring to the conversation |
The common thread across the four is a shift from plans to outcomes: from “do you have a business continuity plan” to “which services matter, how long can each be down, and can you prove you would recover inside that time.” That is the question a resilience audit answers, and the Topical Requirement’s twenty requirements are organized to get there. The site’s guide to the IIA Topical Requirements covers the documentation rules common to all of them, and the financial services guide the supervisory context for DORA and the UK regime.
The vocabulary that has to be precise
Continuity audits go wrong in the definitions, because the terms are used loosely by the people who own the plans and an auditor who accepts the loose usage cannot test anything. Business continuity management is the discipline of keeping critical business processes running through disruption, by any means, including manual workarounds. Disaster recovery is the technology subset: restoring systems and data. Crisis management is the decision-making structure during a disruption: who decides, who communicates, who escalates. Operational resilience is the outcome the three are supposed to produce, expressed at the level of the services customers and the market depend on. The measures that connect them are the ones that get confused. The business impact analysis establishes, for each process, the maximum tolerable period of disruption (MTPD), the point at which the harm becomes unacceptable; the recovery time objective (RTO) is the target for restoring the process, which must be shorter than the MTPD; the recovery point objective (RPO) is how much data loss is acceptable, which determines backup frequency; the minimum business continuity objective (MBCO) is the reduced level of service that is acceptable during recovery; and an impact tolerance, in the UK regime’s sense, is the maximum disruption to an important business service the firm will tolerate, set by the board and expressed in time and, where relevant, volume or value. The first test in any resilience audit is whether the organization’s RTOs are shorter than its MTPDs, whether the RPOs match the backup schedules, and whether anyone can explain where the numbers came from; the worked example below turns on exactly that.
The twenty requirements and the evidence for each
The Topical Requirement has six governance requirements, four risk management requirements, and ten control requirements. The table paraphrases each and names the evidence that supports a conclusion; the requirement document is the criteria to cite. Read down the controls column and notice how much of it is technology: data classification and backups, critical IT controls, the IT asset inventory, and disaster recovery testing account for half the control requirements, which is a fair reflection of where disruptions now come from.
| Ref | Requirement (paraphrased) | What effective looks like | Evidence to obtain |
|---|---|---|---|
| Gov A | A formal resilience strategy with operational, technological, and financial elements, approved by the board and aligned with risk management | A board-approved strategy that names the services that matter, the tolerances, and the investment to meet them | The strategy document; board approval minute; alignment to the risk register and appetite |
| Gov B | Periodic board updates on resilience objectives and oversight progress | A fixed cadence of resilience reporting with test results, tolerance breaches, and remediation | The reporting calendar; the last four packs; minutes showing challenge |
| Gov C | Policies and procedures for critical processes, reviewed and updated regularly | Continuity, DR, and crisis policies with owners, review dates, and version control; procedures at process level | Policy inventory; review evidence; sampled procedures for critical processes |
| Gov D | An incident command structure with defined roles, decision hierarchies, and escalation protocols | A crisis management team with named roles, deputies, decision rights, and escalation thresholds, exercised | The structure document; call trees tested; exercise records showing the structure operating |
| Gov E | A process to validate the competencies of personnel in critical resilience roles | Role requirements defined; training and exercise participation tracked; deputies competent | Role descriptions; training records; exercise attendance; a test that a deputy can execute |
| Gov F | Stakeholder identification and engagement in resilience planning and reporting | Customers, regulators, suppliers, and staff identified with communication plans for disruption | Stakeholder map; communication templates; evidence of engagement |
| RM A | Continuous identification, assessment, and management of resilience risks mapped to strategic objectives | Resilience risks in the register with owners, assessed against the services they threaten | Register entries; the mapping to services; assessment refresh dates |
| RM B | A designated individual or team accountable for monitoring and reporting resilience risks | A named owner with the mandate and the data | Role description; reporting produced by the owner |
| RM C | A process to monitor and escalate unacceptable resilience risks against established tolerance levels | Tolerances set; monitoring against them; escalation when breached | Tolerance definitions; monitoring reports; escalation records |
| RM D | An implemented and periodically tested incident response and recovery process covering detection, containment, recovery, and post-incident analysis | The full lifecycle exists and has been exercised, with lessons recorded | Process documentation; exercise reports; incident post-mortems |
| Ctl A | A third-party assessment process identifying critical suppliers and alternative vendors | Critical suppliers identified from the BIA; their resilience assessed; substitutes or exit plans in place | Supplier criticality list; assessments; exit plans; the register of arrangements where DORA applies |
| Ctl B | Data classification identifying critical data and backup and recovery procedures | Critical data classified; backup frequency matches RPO; recovery procedures documented | Classification; backup configuration; RPO reconciliation |
| Ctl C | Critical IT controls and continuous monitoring for cybersecurity and data protection | The cyber controls that protect recoverability (immutable backups, segmentation, privileged access, monitoring) in place and monitored | Control evidence; monitoring output; reference to the cybersecurity engagement |
| Ctl D | An inventory of critical IT assets needed for operational continuity | A current inventory mapped from critical processes to applications, infrastructure, and data | The inventory; reconciliation to the BIA and to the CMDB |
| Ctl E | Business continuity and disaster recovery plans with periodic testing and board reporting | Plans per critical process and system; a test program covering them on a cycle; results reported upward | Plans; test schedule and reports; board reporting |
| Ctl F | A process to modify the working environment during disruption (remote work, alternate locations) | Tested arrangements for staff to work elsewhere with access to the systems they need | Remote access capacity tests; alternate site arrangements; evidence from real disruptions |
| Ctl G | Continuous monitoring of emerging threats, vulnerabilities, and whistleblower activity | Threat intelligence and vulnerability management feeding the resilience risk assessment; speak-up channels monitored for resilience signals | Monitoring outputs; evidence of action on threats |
| Ctl H | Personnel training on resilience policies with scenario simulation exercises | Training for all, exercises for those with roles, scenarios that change | Training records; exercise scenarios and participation |
| Ctl I | Budgeting and availability of operational, human, technological, and financial resources | Resilience funded; resources available at the time of disruption (spare capacity, contracts, liquidity) | Budget; resource plans; liquidity arrangements for the financial dimension |
| Ctl J | A post-incident review and lessons-learned process integrated into future planning | Every incident and exercise produces actions that are tracked and change the plans | Post-incident reports; action log; plan version history showing the changes |
The requirement that separates a real program from a paper one is Ctl E read together with RM D and Ctl J: plans that are tested, incidents that are analyzed, and lessons that change the plans. An organization can satisfy the other seventeen with documents; those three require the cycle to have turned. The Third-Party Topical Requirement guide covers Ctl A in the depth it needs, since critical suppliers are a resilience topic and a third-party topic at once.
Auditing the business impact analysis and the criticality tiers
The business impact analysis is the foundation of everything else, and it is where an auditor should spend the first third of the fieldwork, because a wrong BIA makes every plan built on it wrong. A BIA takes each business process, establishes what happens over time if it stops (financial, customer, regulatory, reputational, safety), derives the MTPD from the point at which the harm becomes unacceptable, sets an RTO inside it, sets an RPO from how much data the process can lose, and maps the process to the applications, data, people, sites, and suppliers it depends on. The audit tests four things. Completeness: does the process inventory cover the organization, or only the head office functions that filled in the template? Basis: are the MTPDs derived from analysis of harm over time, or did each owner write “24 hours” because that is what the last owner wrote? Consistency: are the RTOs shorter than the MTPDs, are the RPOs achievable with the actual backup schedule, and do dependent processes have compatible targets, since a process with a four-hour RTO that depends on a system with a 72-hour RTO has a 72-hour RTO? Currency: has the BIA been refreshed after the reorganization, the acquisition, the new ERP module, the move to the cloud? A tier model like the one below is what a mature BIA produces, and its absence is usually the first finding.
| Tier | Definition | Typical MTPD | RTO | RPO | Recovery strategy | Examples |
|---|---|---|---|---|---|---|
| Tier 0: continuous | Disruption causes immediate, unacceptable harm; the organization cannot operate without it | Under 4 hours | Under 1 hour | Near zero | Active-active or hot standby; automatic failover; manual workaround for minutes only | Payment processing at a bank; order capture at a retailer; dispatch at a distributor |
| Tier 1: critical | Harm becomes severe within a day | 4 to 24 hours | 2 to 8 hours | Under 1 hour | Warm standby; frequent replication; documented manual procedures for the gap | Route settlement and cash deposit; warehouse management; treasury payments; patient records |
| Tier 2: important | Harm becomes serious within days | 1 to 3 days | 8 to 24 hours | 4 to 24 hours | Restore from backup to alternate infrastructure; manual workaround for a day or two | Accounts payable; payroll (outside the pay run window); HR systems; reporting |
| Tier 3: deferrable | Can stop for a week or more with manageable harm | Over 3 days | 1 to 5 days | 24 hours or more | Standard restore; rebuild if necessary | Training platforms; internal collaboration tools; non-urgent analytics |
Three tests on the tiers are cheap and decisive. Count the Tier 0 and Tier 1 processes; if more than a fifth of the inventory is Tier 0, the BIA was completed by owners rating their own importance, and the recovery investment implied is unaffordable, which means the real tiers are being set by the IT budget rather than the analysis. Trace three Tier 1 processes to their systems and check the systems’ RTOs against the processes’ MTPDs; a mismatch is the single most common critical finding in this domain. And ask who signed the BIA at the top; a BIA the board has never seen fails Gov A. The operational risk guide covers how the BIA feeds the wider operational risk assessment.
Testing: the hierarchy, a credible program, and how tests fail
A plan is a hypothesis until it is tested, and tests come in a hierarchy of realism and cost. The credible program uses all levels on a cycle, with the most realistic tests reserved for the highest tiers, and it treats a test that finds nothing as a test that was too easy. The table gives the hierarchy, what each level can and cannot prove, and the frequency a Tier 1 process should see.
| Test type | What it involves | What it proves | What it cannot prove | Frequency for Tier 1 |
|---|---|---|---|---|
| Plan review and call-tree test | Read-through of the plan; contact every person in the tree and time the response | The plan is current; people can be reached | Anything about capability | Twice a year |
| Tabletop exercise | The crisis team walks a scenario in a room, making decisions in sequence | Decision structure works; roles are understood; gaps in the plan surface | That systems recover or staff can work | Annually, with a new scenario each time |
| Technical component test | Restore a backup; fail over one system; bring up the alternate network | A specific technical capability, with a measured time | The end-to-end process | Quarterly restores for Tier 0 and 1 data; annual failover per system |
| Simulation with business participation | A process is run on recovered systems or manual procedures by the people who actually perform it, against test transactions | The process can operate in recovery mode; the workaround works; the RTO is realistic | Behavior under real stress; concurrent failures | Annually per Tier 1 process |
| Full failover or live test | Production moved to the recovery environment, or a site evacuated, for real | End-to-end recovery within RTO under near-real conditions | The specific scenario that actually occurs | Every one to two years for Tier 0 systems; where the risk of the test itself is manageable |
| Third-party joint test | Critical supplier participates in a scenario or demonstrates its own recovery | The supplier’s plan is real and the interface works | The supplier’s behavior when every client calls at once | Annually for the most critical suppliers; DORA expects it for critical ICT providers |
| Threat-led or adversarial test | Red team simulates a ransomware attack or destructive event; recovery is executed as if real | Detection, containment, and recovery under an intelligent adversary | Everything not in the scenario | Every three years for the largest financial firms under DORA; as risk warrants elsewhere |
The audit of testing asks five questions. Was the schedule met, and what was skipped? Did the tests measure time, so that RTO achievement is a number rather than an adjective? Did they involve the business, or only IT? Were the scenarios realistic and varied, or the same power outage every year? And what happened to the failures: were they recorded, actioned, and retested, or written up as “lessons learned” and filed? A test program in which every test passes is a finding in itself, because real recovery is hard and a program that never fails is not testing what would fail. The ITGC primer covers the backup and change controls that technical tests depend on.
Where plans break: third parties, technology, backups, and people
Plans fail at their dependencies, and four dependencies account for most failures. Third parties: the plan assumes the payroll provider, the cloud platform, the logistics partner, or the single-source supplier will be there, and the audit asks whether each critical supplier’s own resilience has been assessed, whether the contract gives the organization the rights it needs in a disruption (data return, step-in, notification), and whether a substitute or an exit plan exists; a concentration of critical processes on one provider with no substitute is a finding whatever the provider’s SOC 2 says. Technology: the recovery environment has to be current with production, which fails when changes are deployed to production and not to the standby, when licenses do not cover the recovery site, or when the network path to the recovery site has never carried real load. Backups: the questions are whether backups cover the critical data identified in the BIA at the RPO frequency, whether they are immutable or offline so that a ransomware attacker who reaches production cannot reach them, whether restores are tested (not backups: restores) and timed, and whether the restore time meets the RTO; the 3-2-1 rule (three copies, two media, one off-site) is the floor, and immutability is the 2026 expectation. People: the plan names individuals who have left, assumes a deputy who has never done the job, and relies on knowledge that lives in one head, which is why Gov E requires competency validation and why the audit tests a deputy rather than the primary. The Third-Party Topical Requirement guide covers the supplier assessment in full and the cloud risk guide the shared-responsibility questions for recovery in cloud services.
The work program
The program below assesses all twenty requirements for a mid-sized organization with a dozen Tier 1 processes and is sized at 380 hours; a financial institution under DORA or the UK regime adds the regulatory artifacts and roughly 40 percent. The distinguishing feature is the amount of re-performance: the auditor observes a restore, times it, and follows a recovered process through a simulation, rather than reading the test report.
| Step | Procedure | Requirements | Hours |
|---|---|---|---|
| 1. Governance and strategy | Read the strategy, policies, and a year of board reporting; interview the accountable executive, the crisis lead, the CIO, and the head of risk; assess the incident command structure and stakeholder plan | Gov A to F, RM B | 40 |
| 2. Risk management | Assess resilience risks in the register, the tolerances, the monitoring and escalation process, and the incident lifecycle; review the last three real incidents end to end | RM A, RM C, RM D, Ctl G, Ctl J | 40 |
| 3. BIA completeness and basis | Reconcile the process inventory to the organization chart and the application inventory; test the derivation of MTPD, RTO, and RPO for ten processes; check currency against the change log | Ctl D, Ctl E (basis) | 50 |
| 4. Dependency mapping | For every Tier 0 and 1 process, trace to applications, data, sites, suppliers, and people; compare system RTOs to process MTPDs; identify single points of failure | Ctl A, Ctl D | 40 |
| 5. Plans | Sample plans for six Tier 1 processes and four Tier 1 systems: completeness, currency, manual workarounds, named roles, deputies | Gov C, Gov D, Gov E, Ctl E, Ctl F | 30 |
| 6. Backup and recovery re-performance | Reconcile backup configuration to RPOs and classification; observe and time a restore of Tier 1 data; verify immutability and off-site copies; review the last four quarters of restore tests | Ctl B, Ctl C | 40 |
| 7. Test program | Assess the schedule against the cycle policy; review every test report for the year; verify failures were actioned and retested; attend or observe one simulation | Ctl E, Ctl H, Ctl J | 40 |
| 8. Third parties | For the ten most critical suppliers: assessment on file, contract rights, substitute or exit plan, joint test evidence; where DORA applies, the register of information | Ctl A | 30 |
| 9. Scenario walkthrough | Run a ransomware or site-loss scenario with the crisis team and two process owners, unannounced as to timing; observe decisions, timings, and gaps | Gov D, RM D, Ctl F, Ctl H | 24 |
| 10. Resources | Assess the resilience budget, spare capacity, and the financial dimension (liquidity and insurance arrangements for a prolonged disruption) | Ctl I | 16 |
| 11. Conclusions and reporting | Twenty-requirement assessment table with rationale for any exclusion; findings; report | All | 30 |
| Total | — | — | 380 |
Worked example: MidState Beverage’s route accounting module
MidState Beverage is a three-state drinks distributor with twelve depots, 300 delivery routes, and about $31 million a year in cash and checks collected by drivers and settled daily through a route accounting module bolted onto the ERP in 2013, with handheld devices syncing each route’s sales and collections at end of day. Its FY28 audit plan included a resilience engagement scoped to the Topical Requirement ahead of the April 2027 effective date, and the route accounting module became the case study, because it is the process the business cannot run without and the one nobody had thought of as critical.
| Area | What the engagement found | Requirement | Rating |
|---|---|---|---|
| BIA and tiering | The BIA, last refreshed in FY24, rated the route accounting module Tier 2 with a 72-hour RTO, set by IT from the age of the platform. Walking the process showed that if the module is down, drivers cannot settle, cash and checks accumulate in trucks and depot safes at about $120,000 a day across twelve depots, customers cannot be invoiced, and after two days routes cannot be loaded because inventory positions are unknown. The MTPD derived from that analysis was about eight hours; the actual recovery target was nine times longer than the process could tolerate | Ctl D, Ctl E; Gov A | High |
| Manual workaround | The plan’s workaround was “depots settle on paper and key later.” At the two depots tested, no paper settlement forms existed, no one under 40 had ever done a paper settlement, and the module could not accept back-dated entries without an administrator override, which was the control the FY27-01 route cash audit had just restricted | Ctl E, Ctl F, Gov E | High |
| Backups and restore | Nightly backups met a 24-hour RPO that the BIA had never questioned; a day of settlements is about $120,000 of cash movements, which the business judged unacceptable once asked. The last restore test of the module was in FY23. The engagement observed a restore to a test environment: it succeeded in 11 hours, against the 72-hour RTO and the 8-hour MTPD. Backups were online and writable from the production domain | Ctl B, Ctl C | High |
| Dependencies | The handheld sync service ran on one server in the head office with no standby; the vendor that maintained the 2013 module was a two-person firm with no continuity plan and no escrow of the source code; the two acquired distributors ran separate settlement systems that nobody in IT could restore | Ctl A, Ctl D | High |
| Governance and testing | The board had received one resilience update in three years; no exercise had included operations; the crisis plan named a COO who left in FY25 | Gov B, Gov D, Ctl H, Ctl J | Medium |
| What worked | The ERP core and finance systems had a warm standby tested annually with measured four-hour recovery; the cybersecurity program’s segmentation would have limited a ransomware event to one domain | Ctl C, Ctl E | Satisfactory for those systems |
| Overall | Unsatisfactory: the organization’s most operationally critical process had been classified from the age of its software rather than from its impact, and every downstream decision (RPO, restore testing, standby, vendor escrow, workaround) had inherited the error | — | Nine actions; $180,000 approved for a standby sync service, a re-tiered backup schedule, immutable backups, source code escrow, and quarterly restore tests, with the ERP migration for the acquired entities brought forward |
The engagement’s value came from one hour of fieldwork: walking the settlement process with a depot manager and asking what happens at 5 p.m. when the module is down. Everything else in the table followed from the answer. A document-based audit would have confirmed that the BIA existed, the plan was dated this year, and a tabletop had been held, and would have rated the area Satisfactory; the Topical Requirement’s insistence on the BIA basis, the dependency inventory, tested recovery, and competent deputies is what turned the same facts into an Unsatisfactory report with $180,000 of remediation the board approved without argument. The walkthrough guide covers the technique that found it.
Metrics for the board
Gov B requires periodic board reporting, and the report is only useful if it carries measures that move. The set below is what a mature program reports quarterly; a board that sees these eight numbers can tell whether resilience is improving without reading a plan.
| Metric | Definition | Target | What a bad number means |
|---|---|---|---|
| BIA currency | Share of Tier 0 and 1 processes with a BIA refreshed in the last 12 months and after any material change | 100 percent | Plans built on an old picture of the business |
| RTO coverage | Share of Tier 0 and 1 processes whose supporting systems have an RTO at or inside the process MTPD | 100 percent | Recovery targets the business cannot survive |
| Restore test success | Share of scheduled Tier 0 and 1 restore tests completed with measured time inside RTO | Above 90 percent completed; 100 percent inside RTO or actioned | Untested or slow recovery |
| Exercise completion | Tests completed against the annual program, by type | 100 percent of scheduled, with business participation in simulations | A program on paper |
| Open test findings | Failures from tests and incidents open beyond their due date | Zero overdue High | Lessons recorded, not learned |
| Critical supplier coverage | Share of critical suppliers with a current resilience assessment and an exit or substitute plan | 100 percent | Concentration with no fallback |
| Backup immutability and off-site coverage | Share of Tier 0 and 1 data with immutable or offline copies at the RPO frequency | 100 percent | Ransomware can destroy the recovery |
| Role readiness | Share of crisis and recovery roles with a trained, exercised deputy | 100 percent | Single points of human failure |
Common mistakes in continuity audits
| Failure | What it looks like | Why it matters | Fix |
|---|---|---|---|
| Auditing documents | Plan exists, dated this year, tabletop held: Satisfactory | None of it proves recovery | Observe a restore; time it; follow a process in recovery mode |
| Accepting the BIA | RTOs taken from the BIA without testing their basis | The most common critical error inherits into every plan | Derive MTPD from harm over time for a sample of processes and compare |
| IT-only scope | Disaster recovery audited; business continuity and crisis management not | Systems recover; the business does not know how to use them | Scope to processes and services, with IT as a dependency |
| Trusting backups | Backup jobs succeed; restores never tested | A backup that cannot be restored in time is not a control | Restore tests, timed, on a cycle, for critical data |
| Ignoring immutability | Online backups writable from production | Ransomware encrypts the recovery along with the data | Immutable or offline copies for Tier 0 and 1 data |
| Tier inflation | Half the inventory rated Tier 0 | Unaffordable targets; real priorities set by budget instead of analysis | Challenge the tiers; a Tier 0 process needs a Tier 0 investment |
| Supplier blind spot | Critical suppliers assumed resilient because they are large | Concentration and no exit plan | Assess, contract, substitute, and test jointly |
| Same scenario every year | The annual power-outage tabletop | The plan is tested for the scenario it was written for | Rotate scenarios; include ransomware, supplier failure, site loss, key-person loss |
| Passing tests | Every test succeeds | The tests are too easy or the results are edited | Treat a clean year as a finding about the program |
| No exclusion record | Requirements not assessed with no rationale | Nonconformance with the Topical Requirement from 30 April 2027 | A twenty-row applicability record in every resilience engagement |
Resilience is the one control domain where the audit can, and should, make the thing happen: restore the backup, run the process, call the deputy, and time it. An organization whose plans have survived that kind of audit will survive most disruptions; one whose plans have only survived a document review has a binder. The disaster recovery guide covers internal audit’s role during an actual disruption, and the risk appetite guide covers how a board sets the tolerances the whole program is built to.
Related guides
- IIA Topical Requirements — the framework and the documentation rules
- Third-Party Topical Requirement — critical supplier assessment in depth
- Auditing cybersecurity programs — the controls that protect recoverability
- Cloud computing audit — recovery under shared responsibility
- ITGC primer — backup and change controls
- Operational risk — where the BIA feeds the wider assessment
- Internal audit and disaster recovery — the auditor’s role during a real event
- Internal audit in financial services — DORA and the UK regime in supervisory context
- Audit walkthroughs — the technique that finds the real MTPD
- Risk appetite statements — setting tolerances at the board
- All Guides — the full index
Leave a Reply