,

Assessing Cybersecurity With NIST CSF 2.0: An Internal Audit Method

NIST’s Cybersecurity Framework is the most common set of criteria internal auditors bring to a cybersecurity assessment, and the most commonly misused. Version 2.0, published on 26 February 2024, added a Govern function, reorganized the outcomes into six functions, 22 categories and 106 subcategories, and kept the two things that make it work as audit criteria: outcome statements that describe a result rather than a control, and a profile mechanism for stating where the organization is and where it intends to be. This guide is an internal audit method for using it: how to scope a first assessment, how to turn subcategories into evaluation criteria with an evidence standard, how to score without inventing a maturity model NIST never published, and how to report a posture view a board can act on.

In this guide

What CSF 2.0 is, in the terms an auditor needs

The framework has three parts. The Core is the taxonomy of outcomes: six functions, each split into categories, each split into subcategories that state an outcome such as “assets are managed” or “adverse events are analyzed”. The Core is deliberately not a list of controls; the same subcategory can be achieved by different controls in different organizations, which is why it works as criteria across industries and why it cannot be tested by ticking a box. Profiles are the mechanism for applying the Core: a Current Profile records which outcomes the organization achieves today and to what extent, a Target Profile records the outcomes it intends to achieve given its risk appetite, obligations and resources, and the gap between the two is the improvement plan. NIST also publishes Community Profiles, sector or use-case baselines built with external contributors, that an organization can adopt as its Target. Tiers, from Tier 1 Partial to Tier 4 Adaptive, characterize the rigor of the organization’s cybersecurity risk governance and risk management practices, as a whole, and are used to inform the Target Profile; they are not a score for each subcategory.

FunctionWhat it coversCategories
Govern (GV)Strategy, expectations and policy for managing cybersecurity risk; new in 2.0 and cross-cuttingOrganizational Context (GV.OC); Risk Management Strategy (GV.RM); Roles, Responsibilities, and Authorities (GV.RR); Policy (GV.PO); Oversight (GV.OV); Cybersecurity Supply Chain Risk Management (GV.SC)
Identify (ID)Understanding the assets, suppliers and risks, and finding improvementsAsset Management (ID.AM); Risk Assessment (ID.RA); Improvement (ID.IM)
Protect (PR)Safeguards that manage the riskIdentity Management, Authentication, and Access Control (PR.AA); Awareness and Training (PR.AT); Data Security (PR.DS); Platform Security (PR.PS); Technology Infrastructure Resilience (PR.IR)
Detect (DE)Finding and analyzing possible attacks and compromisesContinuous Monitoring (DE.CM); Adverse Event Analysis (DE.AE)
Respond (RS)Acting on a detected incidentIncident Management (RS.MA); Incident Analysis (RS.AN); Incident Response Reporting and Communication (RS.CO); Incident Mitigation (RS.MI)
Recover (RC)Restoring assets and operationsIncident Recovery Plan Execution (RC.RP); Incident Recovery Communication (RC.CO)

Two supporting pieces matter for audit use. Informative References map each subcategory to the control catalogs that implement it, including NIST SP 800-53 Rev. 5, ISO/IEC 27001:2022 and CIS Controls, so an assessment can borrow test steps from whichever catalog the organization runs on. And Implementation Examples, published alongside 2.0, give concise, non-mandatory examples of what achieving a subcategory can look like; they are the closest thing NIST offers to a test procedure, and they are explicitly not requirements. NIST’s Quick Start Guides cover the same ground for specific audiences, including small businesses and the creation of profiles.

How assessments go wrong

The first failure is treating the Core as a compliance checklist and asking management, subcategory by subcategory, “do you do this?” The answer is yes to everything, the assessment produces a wall of green, and the organization’s first ransomware event happens in a green category. The second is inventing a maturity model: scoring every subcategory from 1 to 5 on a scale nobody defined, averaging the scores by function and presenting “3.2” to the board as if it measured something. The Tiers exist and are defined, but they describe governance and risk management rigor for the organization, not achievement of each outcome, and NIST never published a per-subcategory maturity scale. The third is assessing without an evidence standard, so that a documented policy, an implemented control and a control that has been shown to operate over a period all earn the same mark. The fourth is scoping the whole Core for a first assessment at an organization that has never had one, which produces 106 conclusions of uneven quality and a report that arrives too late to matter. The method below is built to avoid all four.

Scoping a first-time assessment

A first assessment at an organization with no prior profile has two jobs: give the board a defensible posture view, and leave behind a Current Profile that the next assessment can be measured against. It cannot do both across all 106 subcategories in a reasonable budget with reasonable quality, so scope by depth rather than by omission. Assess every subcategory in Govern and Identify at full depth, because those functions are where a first assessment finds its most important gaps and where the evidence is documents, minutes and registers that a small team can read quickly. Assess Protect, Detect, Respond and Recover at category level, with a sample of subcategories in each category taken to full depth, chosen by the risk assessment and by what the organization has recently suffered or nearly suffered. Record the sampled subcategories and the reason for each choice, so that the untested subcategories are visibly unassessed rather than silently assumed.

Three scoping decisions need to be written down before fieldwork. The organizational boundary: which legal entities, sites, business units and technology environments the profile covers, and the treatment of acquired units, operational technology and outsourced services, each of which is a place a first assessment tends to stop without saying so. The Target: whether the organization has a Target Profile (rare, the first time), whether a Community Profile fits (NIST’s profiles and sector baselines are candidates), or whether the assessment will propose a Target as an output, which is usually the honest answer. And the relationship to other work: the cybersecurity program audit that tests controls in depth, the incident response audit, the backup and recovery audit and the patch and vulnerability audit each supply evidence for whole categories, and the assessment should consume their results rather than re-perform them, with the reliance evaluated and documented.

Turning subcategories into criteria

A subcategory is an outcome statement, and an outcome cannot be tested until the auditor has written down what achieving it would look like here. The conversion has three parts for each subcategory in scope. The interpretation states the outcome in the organization’s own terms: for ID.AM-01, “inventories of hardware managed by the organization are maintained”, the interpretation at a distributor might read “every server, endpoint, handheld and network device that touches the corporate network appears in the configuration database with an owner, and the database is reconciled to independent sources on a cycle”. The criteria are the specific, observable conditions that would demonstrate the interpretation, drawn from the Implementation Examples and from the Informative References for the catalog the organization uses: reconciliation performed, unknown devices investigated, disposal recorded. And the evidence list names what the auditor will obtain and what test will be applied to it. Written this way, the criteria are agreed with management before testing, which removes the argument at the end about what the subcategory meant.

The interpretation step is also where the Target Profile is born. If management cannot say what achieving a subcategory would look like for the organization, the organization has no target for it, and that absence is itself a Govern finding (GV.RM and GV.OC) rather than a reason to skip the row. The risk appetite statements guide covers how to elicit tolerances that can be turned into targets, and the risk register guide the register that a cyber risk assessment feeds.

The evidence standard per subcategory

The single most important discipline is separating three states of achievement that assessments habitually merge. An outcome is documented when a policy, standard or procedure requires it. It is implemented when the control or process that produces it exists and can be observed at a point in time. It is operating when evidence over a period shows it working, with the exceptions handled. Each state needs a different kind of evidence and supports a different conclusion, and the table below sets the standard by function, with the sampling rule that makes the operating conclusion defensible.

FunctionExample subcategoryDocumented (evidence)Implemented (evidence)Operating (evidence and sample rule)
GovernGV.OV-01: cybersecurity risk management strategy outcomes are reviewed to inform and adjust strategy and directionCharter or policy assigning the reviewA review has occurred: minutes, reports presentedReviews across the period on the stated cadence with decisions and follow-through; all instances in the period (usually four or fewer)
IdentifyID.AM-01: hardware inventories are maintainedAsset management standard requiring an inventoryInventory exists with owners and attributesReconciliation to directory, EDR and network sources performed on cycle with unknowns cleared; re-perform one reconciliation and inspect the last three
ProtectPR.AA-05: access permissions, entitlements, and authorizations are defined in a policy, managed, enforced, and reviewedAccess policy and role definitionsAccess reviews configured; provisioning workflow in placeReviews completed for the period with revocations actioned; sample reviews and trace 25 to 40 access changes through approval to system state
DetectDE.CM-01: networks and network services are monitored to find potentially adverse eventsLogging and monitoring standard naming sources and retentionLog sources connected; alert rules definedAlerts triaged within SLA across the period; sample 25 alerts across severities and reconcile log sources to the systems in scope
RespondRS.MA-01: the incident response plan is executed in coordination with relevant third parties once an incident is declaredIncident response plan with third-party contactsPlan current; roles trained; contacts verifiedIncidents in the period handled per plan, or an exercise executed with the third parties; inspect every declared incident and the last exercise
RecoverRC.RP-01: the recovery portion of the incident response plan is executed once initiated from the incident response processRecovery plan with sequence and objectivesBackups configured to the plan; restore procedures documentedRestore tests executed within the period meeting objectives; observe one restore of a crown-jewel system and inspect the last two test records

The rule that follows from the table: a subcategory cannot be concluded as achieved on documented evidence alone, and an implemented conclusion has to say so in the words “implemented, operation not tested”. Where the assessment consumes another engagement’s testing, the operating evidence is that engagement’s, and the record says which and when. Where a subcategory is achieved through a third party, the evidence is the third party’s report plus the organization’s own complementary controls, per the SOC 2 review guide, not the contract.

Fieldwork mechanics: the workshop, the evidence request and the sampling plan

The assessment runs in three passes, and the order matters. The first pass is a structured workshop, half a day per function for Govern and Identify and a day for the other four together, in which the auditor walks the in-scope subcategories with the control owners, records management’s own statement of the current state, and agrees the interpretation and criteria. The workshop output is a table of claims, not conclusions, and it is the most efficient way to find the outcomes nobody owns: the subcategory where the IT manager looks at the plant engineer and both wait for the other to answer is a finding before any evidence is requested. The second pass is the evidence request, generated from the criteria: one request list per function, each item tied to a subcategory and to the evidence state it supports, issued with a date and tracked. Requests that come back empty are recorded as such, because “no evidence provided” and “not achieved” are different conclusions and the report has to say which applies.

The third pass is testing to the evidence standard, and the sampling plan decides how much of the budget it consumes. For period-based outcomes, sample sizes follow the function’s normal attribute-sampling rules by control frequency; the sample sizes guide sets those out. For population-based outcomes, the test is a reconciliation rather than a sample: the asset inventory against independent sources, the log sources against the systems in scope, the supplier inventory against accounts payable, the access list against HR. Reconciliations are where first assessments find their largest numbers, and they cost less than sampling because the data usually exists. The plan states, for every subcategory in scope, which of the three approaches applies (inspection of all instances, attribute sample, or reconciliation) and the expected hours, and the total is compared to the budget before fieldwork starts rather than discovered in week four.

Scoring discipline: outcomes, profiles and tiers

Score each subcategory on achievement of the outcome, not on the sophistication of the control, using a four-value scale whose values are defined by the evidence state above. Then build the Current Profile from the scores, compare it to the Target, and assess the Tier once, at the organization level, from the Govern and Identify evidence. Never average subcategory scores into a function score; report the distribution instead, because an average of “fully achieved” and “not achieved” is a number that describes nothing. The scale, with what earns each value, is the part of the method to agree with management and the audit committee before fieldwork.

ValueDefinitionEvidence requiredWhat it is not
AchievedThe outcome is produced consistently across the period for the whole scopeOperating evidence with sample results within tolerance; scope coverage confirmed (no excluded units or systems)Not a statement that the control is best practice or that the Tier is high
Largely achievedThe outcome is produced, with defined exceptions in scope or period that are known and managedOperating evidence with exceptions listed, owned and time-bound; the excluded scope namedNot “implemented but untested”; that is the next value
Partially achievedThe outcome is documented and implemented but operation is not evidenced, or operation fails for a material part of the scopeImplemented evidence; either no period evidence or sample failures beyond toleranceNot a placeholder for “we are working on it”
Not achievedThe outcome is not documented, not implemented, or the process exists only on paperAbsence of implemented evidence, or documented onlyNot a judgment of intent or effort

The Tier assessment is separate and short. Tier 1 Partial describes ad hoc, reactive risk management with limited organizational awareness; Tier 2 Risk Informed describes risk-aware practices approved by management but not established organization-wide; Tier 3 Repeatable describes formally approved policy, organization-wide practices and regular updates informed by changes in risk; Tier 4 Adaptive describes practices continuously improved from lessons learned and predictive indicators, with cybersecurity risk management part of the organizational culture. The auditor states the Tier the evidence supports, the Tier the organization’s risk profile and obligations argue for as a target, and the specific Govern subcategories that separate the two. That is the whole Tier discussion; it does not need a score.

Reporting a defensible posture view

The board wants one page, and the one page has four elements. A distribution by function: for each of the six functions, how many subcategories in scope fall in each of the four values, with the unassessed count shown separately, so the reader sees both the posture and the coverage. The Current-to-Target gap: the subcategories where the Target requires “achieved” and the Current is two or more values below, ranked by the risk assessment, which is the improvement plan in embryo. The Tier statement: current, target and the three or four governance conditions that separate them. And the crown-jewel view: for the handful of systems whose loss would stop the business, the Protect, Detect and Recover outcomes specifically for those systems, because an enterprise-wide “largely achieved” can hide a crown jewel that is not backed up. The audit committee presentation template has a layout that fits; the detailed subcategory records go in an appendix that nobody on the board reads and everybody’s counsel keeps.

Findings are written against subcategories, not against functions, and rated on the function’s normal severity scale rather than on the CSF score, per the finding severity ratings guide: a “not achieved” on GV.SC-06 (due diligence on suppliers before contracting) at an organization that outsources its data center is a High finding; the same value on a subcategory covering a low-risk internal process may be a Low. The report says both the CSF value and the severity, and explains the difference where a reader might expect them to align.

One more audience needs the same material in a different shape. Customers, insurers and regulators increasingly ask for a “CSF posture statement”, and management will want to answer from the assessment. The statement they can honestly make is the distribution by function with the scope and depth stated, the Tier the evidence supports, and the date; it is not a certification, the framework has no certification, and an internal audit function should not let its report be quoted as one. The report should say so in a sentence, and the sentence should be the one management copies into the questionnaire.

The subcategory assessment record

One record per subcategory in scope, in the workpaper system. The set of records is the Current Profile; the Target column is what makes the gap visible without a separate document.

CSF 2.0 subcategory assessment record

1. Reference. Function, category and subcategory identifier and NIST outcome text; scope depth (full or sampled) and the reason for selection if sampled.

2. Interpretation. The outcome restated in the organization’s terms, agreed with the control owner, with the boundary (units, sites, systems) to which it applies.

3. Criteria. The observable conditions that demonstrate the interpretation; the Informative References or Implementation Examples drawn on; the catalog controls (for example ISO/IEC 27001:2022 Annex A or CIS Controls) that implement it here.

4. Target. The value the Target Profile requires (or the proposed target if none exists) and its source: risk appetite, regulation, contract, Community Profile.

5. Evidence. Documented evidence obtained; implemented evidence obtained; operating evidence obtained with sample sizes, periods and results; reliance on other engagements with references.

6. Current value. Achieved / largely achieved / partially achieved / not achieved, with the basis sentence naming the evidence and the excluded scope, if any.

7. Gap and finding. Gap to Target in values; finding reference and severity if raised; root cause category (governance, resourcing, technology, third party, process).

8. Sign-off. Assessor, reviewer, control owner acknowledgment of the facts, dates.

Worked example: Brightwater Foods’ first assessment

Brightwater Foods is the illustrative food manufacturer used elsewhere on this site: about $180 million of revenue, 600 staff, three plants, a co-sourced internal audit function that the finance director sponsors. It had never had a cybersecurity assessment. The board asked for one after a customer’s supplier questionnaire required a “NIST CSF posture statement” and nobody could produce one. The engagement was budgeted at 240 co-sourced hours plus 60 internal, and scoped as the method describes: all 31 Govern and all 21 Identify subcategories at full depth, and 24 sampled subcategories across the other four functions (eight in Protect, five in Detect, six in Respond, five in Recover), chosen from a two-day risk workshop and the two events the company had already lived through, a business email compromise that diverted a supplier payment and a plant line stoppage after a contractor’s laptop introduced malware to the historian network. Thirty subcategories were left unassessed and listed as such.

The organizational boundary was the three plants, the head office and the two cloud tenants, with operational technology at the plants explicitly inside the scope, because leaving it out would have made the assessment describe the part of Brightwater that cannot stop production. No Target Profile existed; the assessment proposed one as an output, anchored on the customer’s requirements, the cyber insurer’s conditions and the board’s stated tolerance that a plant outage of more than 48 hours was unacceptable. The interpretation and criteria for each subcategory were agreed with the IT manager and the plant engineering lead before testing, which took most of the first week and saved the last two.

Function (subcategories assessed)AchievedLargely achievedPartially achievedNot achievedUnassessed
Govern (31 of 31)3511120
Identify (21 of 21)34860
Protect (8 of 22)232114
Detect (5 of 11)11216
Respond (6 of 13)01327
Recover (5 of 8)01223
Total (76 of 106)915282430

The distribution told the story before any finding did. Twelve of the 31 Govern outcomes were not achieved, and the twelve clustered in GV.RM and GV.SC: there was no cyber risk management strategy, no risk appetite for cyber, no supplier due diligence before contracting and no supplier inventory that distinguished the co-packer with remote access to the plant network from the stationery vendor. Identify was better on assets (the IT manager kept a good inventory of what IT owned) and poor on risk (ID.RA had no method, no register and no re-assessment after the two incidents). The sampled Protect subcategories were the strongest area because the IT manager had built them, and the sampled Recover subcategories the weakest because nobody had ever restored the ERP or the plant historian end to end. The Tier assessment was one paragraph: Tier 1 Partial on the evidence, with management’s practices reactive and the board’s awareness limited to the questionnaire that had triggered the engagement; a target of Tier 3 Repeatable within 24 months, contingent on the four Govern conditions the report named: an approved risk management strategy with appetite, an accountable owner reporting to the board twice a year, a supplier risk process, and a policy set refreshed on a cycle.

Nine findings were raised, three High: no cybersecurity risk assessment or appetite (GV.RM-01, GV.RM-02, ID.RA-01, ID.RA-05); plant operational technology networks flat with the corporate network and reachable by contractors over shared remote-access credentials (PR.AA-01, PR.AA-03, PR.IR-01); and recovery never demonstrated for the ERP or the plant historian, whose backups sat on the same network segment as the systems they protected (PR.DS-11, RC.RP-01, RC.RP-02). The board page carried the distribution table, the Current-to-Target gap ranked by risk (the three High findings plus GV.SC-06 supplier due diligence and DE.CM-01 monitoring of the plant networks, which had none), the Tier statement, and the crown-jewel view for the ERP, the three plant control networks and the customer EDI gateway. The finance director’s cover note said what the number-averaging approach would have hidden: the company’s technical defenses were better than its governance, and the governance gaps were the reason the technical defenses stopped at the plant door.

Where the assessment meets the Topical Requirement

Since 5 February 2026, an assurance engagement on cybersecurity must apply the IIA’s Cybersecurity Topical Requirement: seventeen requirements across governance, risk management and control processes, each assessed or documented as not applicable. A CSF 2.0 assessment scoped as above satisfies most of the requirement’s rows by construction, because the Govern and Identify functions cover the requirement’s first two domains and the sampled Protect, Detect, Respond and Recover subcategories supply evidence for the third. The efficient sequence is to run the CSF assessment, then complete the requirement’s seventeen-row conformance record from the subcategory records, using the crosswalk in the Topical Requirement workbook, and to report the seventeen rows as an appendix alongside the CSF posture page. Two requirement rows usually need evidence the CSF sample does not supply, the talent management process for security operations and the endpoint communication controls, so the sampling plan should include a Protect subcategory for each. The requirement is mandatory for assurance and recommended for advisory; a first CSF assessment framed as advisory can still be documented to the requirement, and usually should be, because the next one will be assurance.

Repeat assessments: measuring movement between profiles

The second assessment is where the method pays for itself, provided the first one left behind subcategory records with agreed interpretations. The repeat holds the interpretations and criteria constant unless the environment changed, re-tests the subcategories that were “partially” or “not achieved” against their remediation, extends depth into the subcategories the first assessment left unassessed, and reports movement as counts by value per function rather than as a change in an average. Three reporting rules keep the movement honest. A subcategory that moved from “not achieved” to “achieved” in one cycle needs operating evidence over the period, not a completed project; if the control was implemented in the last quarter, the value is “partially achieved” with the implementation date, and it moves next time. A subcategory whose interpretation changed is reported as re-baselined, not improved. And the unassessed count has to fall each cycle, so that by the third assessment the Current Profile covers the whole Core at the depth the risk warrants; a program that samples the same 24 subcategories every year is measuring its comfort, not its posture.

The Target Profile also moves. Obligations change (a customer contract, a regulator, an insurer), the organization changes (an acquisition adds a plant, a cloud migration removes a data center), and the risk appetite the board stated in year one is tested by what year one cost. The repeat assessment re-confirms the Target with the sponsor before fieldwork and records the changes, so that the gap being reported is against a target somebody currently owns. That is also the point at which the Tier target is re-examined: an organization that has closed the Govern gaps is usually ready to state a Tier 3 target with a straight face, and one that has not should hear that from the auditor before it hears it from a customer.

Common mistakes

Assessing all 106 subcategories at a first engagement, thinly. Scoring on a home-made 1 to 5 maturity scale and averaging it, then presenting the average as posture. Concluding “achieved” on documented evidence, or on management’s assertion in a workshop. Leaving operational technology, acquired units, cloud tenants and third parties outside the boundary without saying so on the report’s first page. Treating the Tiers as per-subcategory scores. Producing a Current Profile with no Target, so the gap and the plan never appear. Writing findings against functions (“Recover is weak”) instead of subcategories with evidence. Re-performing testing that the incident response, backup and vulnerability audits already did, instead of relying on it with the reliance documented. And issuing the report without leaving behind the subcategory records that make the next assessment a comparison rather than a fresh start; the IT audit plan guide shows where the follow-on assessment sits in the cycle, and the ITGC primer covers the general controls layer that the Protect function assumes.

Comments

Leave a Reply

Discover more from internalauditguide.com

Subscribe now to keep reading and get access to the full archive.

Continue reading