,

Vendor and Third-Party Models: Validating What You Cannot See Inside

A growing share of the models that drive decisions in any organisation were built by someone else. The credit scorecard inside the loan origination system, the fraud score on every card transaction, the allowance engine that computes expected credit losses, the pricing tool, the anti-money-laundering transaction monitor, the demand forecaster in the planning suite, and now the language model behind the customer chatbot: each is a model, each carries model risk, and each arrives as a black box whose code, training data and assumptions the buyer cannot inspect. The regulators’ position has been consistent since the 2011 model risk guidance and survives into its 2026 replacement: vendor does not mean exempt. The organisation that uses the model owns its risk, and it has to validate what it cannot see inside. This guide sets out how: the types of vendor model and what is visible in each, the due-diligence substitutes for transparency, the boundary at which your configuration becomes your model, compensating monitoring when validation is limited, the contractual protections that make the rest possible, inventory and contingency, a validation program, and a worked example from a regional bank with a vendor allowance engine and a vendor fraud score.

It extends the model risk and algorithm audit guide, which covers the framework as a whole, into the case that framework handles worst, and it is the model-risk chapter of the third-party cluster that begins with the third-party risk management program guide. Where regulation is cited it is the US model risk guidance and its April 2026 successor, the PRA’s model risk management principles (SS1/23), and, for AI systems, the EU AI Act as amended by the 2026 Digital Omnibus, which deferred the high-risk obligations to 2 December 2027.

In this guide

The black-box problem, and why vendor does not mean exempt

Model validation, as every framework describes it, has three parts: evaluation of conceptual soundness, ongoing monitoring, and outcomes analysis. For a model the organisation built, all three are available: the developers can be interviewed, the code read, the data traced, the assumptions challenged. For a vendor model the first part is largely closed. The vendor treats the methodology as intellectual property, the training data as confidential, and the code as the product; the buyer receives a user manual, a marketing white paper and an output. The temptation is to treat the vendor’s reputation as the validation, and the 2011 US guidance was written partly to end that practice: it required vendor models to be brought into the organisation’s model risk framework, required the organisation to obtain developmental evidence from the vendor, to validate the model with the same rigour it would apply to its own, to understand the limits of what it could validate and compensate for them, and to plan for the model’s termination. The guidance that replaced it in April 2026 (SR 26-2 and OCC Bulletin 2026-13) does not relax that expectation, and the PRA’s principles apply it to every model in the inventory regardless of origin.

The reason the expectation holds is practical rather than legal. The vendor calibrated its model on other customers’ data, for the average of its customer base, at a point in time. The organisation applies it to its own portfolio, its own customers, its own controls, and its own regulators, and the differences between the average customer and this customer are exactly where a vendor model fails: a fraud score tuned for large issuers over-alerts a regional bank’s small-ticket portfolio, an allowance engine’s default assumptions do not match a lender’s concentrated book, a demand forecaster trained on retail patterns misreads a distributor’s route economics. Nobody at the vendor knows those differences; only the user can find them, and only by validating the model against its own outcomes. That is the whole discipline of this guide: where you cannot see inside, look harder at what comes out.

Types of vendor model, and what you can actually see in each

“Vendor model” covers very different things, and the validation approach depends on which one is in front of you. The table sorts them by how much the buyer can see and what lever it has, because the lever decides the program.

TypeExamplesWhat you can seeYour validation lever
Configurable engine with a published methodologyAllowance and expected-loss platforms, asset-liability and interest-rate models, capital and stress platforms, actuarial enginesThe methodology, often in a documented technical manual; the organisation’s own inputs, assumptions and overlays; full outputs and intermediate calculationsConceptual soundness of the methodology can be assessed; the organisation’s configuration is the organisation’s model and is fully validatable; replication of key calculations is often possible
Proprietary score or ratingBureau credit scores, vendor fraud scores, application scorecards embedded in origination systems, market-risk factor modelsInputs the organisation supplies, the score, the vendor’s documentation on intended use and population; sometimes reason codes; rarely the weightsOutcomes analysis on the organisation’s own population; benchmarking against an internal or alternative model; population stability; cut-off and threshold decisions, which are the organisation’s
Rules-and-learning hybridTransaction monitoring for financial crime, fraud detection with tuned rules plus machine-learned componentsThe rules and thresholds the organisation set; the vendor’s learned component only through its alertsRule and threshold validation is fully in scope; the learned component is validated through alert outcomes, false-positive and detection rates, and coverage testing against known typologies
Embedded analytic in a business applicationDemand forecasting in planning suites, pricing optimisation, inventory replenishment, workforce scheduling, marketing propensityInputs, outputs, some parameters; the method often undocumented beyond a help pageOutcomes against actuals; challenger comparisons; the decision process that consumes the output, which is where control usually has to sit
Foundation and generative models via APILanguage models behind assistants, document extraction, summarisation and classification servicesNothing of the model; the prompts, retrieval data and guardrails the organisation built; outputsEvaluation sets built from the organisation’s own use cases; output monitoring; human-in-the-loop controls; the organisation’s layer is the model it owns; see AI prompts for internal auditors for the usage side

The first type is the most common in finance and the least difficult, because its methodology is usually documented and its calculations replicable; the organisation’s problem there is its own configuration, covered below. The last type is the newest and the one where the framework is still catching up; for it, the honest position is that the vendor’s model cannot be validated by the user at all, and the organisation’s controls have to sit entirely in the layer it built and in what it does with the output.

Due-diligence substitutes for transparency

Where the vendor will not open the box, the organisation has six substitutes, and a validation file for a vendor model should show which of them were obtained and which were refused, because a refusal is itself evidence about the model’s risk. Developmental evidence is the first and the one the guidance names: the vendor’s own documentation of the model’s design, components, intended use, the population it was developed on, its limitations and its known weaknesses, usually available under confidentiality to serious customers and usually far more informative than the marketing material. The vendor’s own validation or testing results are the second: reputable vendors validate their models and will share the reports, the performance statistics by segment, and the results of their own outcomes analysis; a vendor with none is a finding. Test-deck rights are the third: the ability to run a set of the organisation’s own cases through the model and compare the results with expectations, with a benchmark model, or with the previous version, which is how most configuration changes and version upgrades should be accepted.

Benchmarking is the fourth: an internal challenger model, a second vendor’s score, or a simple statistical model built on the organisation’s own data, run alongside the vendor model so that divergence can be investigated; for scores, even a coarse internal scorecard reveals where the vendor model ranks customers implausibly. Peer evidence is the fifth: user groups, regulatory findings at other users, published performance comparisons, and the model’s history of methodology changes, which vendors announce to their customers and few customers read. The sixth is the vendor’s independent assurance: for platforms delivered as a service, the assurance report over the platform’s change management and computation controls, read with the method in the SOC 2 review guide, establishes that the model the organisation validated is the model that runs, which the other five cannot. None of the six replaces conceptual soundness review; together they let the organisation write a defensible statement of what it knows about the model, what it does not, and what it did about the gap.

The customisation boundary: when your configuration becomes your model

The most consequential and least understood principle in vendor model risk is that the organisation’s configuration is the organisation’s model. The vendor supplies an engine; the organisation chooses the assumptions, segments, thresholds, overlays, data mappings and overrides that turn the engine into a number, and every one of those choices is a modelling decision the organisation made, must document, and must validate with full rigour, whatever the vendor’s methodology says. Regulators have been explicit that a vendor allowance platform with the organisation’s qualitative factors in it is the organisation’s allowance model, and the same holds for a fraud engine with the organisation’s thresholds and a forecasting suite with the organisation’s seasonality settings. The table draws the boundary for the common configuration elements.

Configuration elementWhose model is it?What the organisation must evidence
Methodology selection (which of the vendor’s methods, segmentation approach)Shared: the vendor’s method, the organisation’s choiceThe rationale for the choice against the portfolio; the alternatives considered; approval by the model owner and the risk function
Input data mapping and transformationsOrganisation’sData lineage, reconciliation of inputs to source, treatment of missing data, change control on mappings; the evidence standard in the IPE testing guide
Assumptions and parameters (loss rates, prepayment speeds, discount rates, macro scenarios, seasonality)Organisation’sSource and support for each assumption, sensitivity analysis, back-testing, governance of changes
Segments and cohortsOrganisation’sSegmentation rationale, homogeneity testing, minimum sizes, treatment of new products
Thresholds, cut-offs and rules (fraud alert thresholds, score cut-offs, monitoring rules)Organisation’s, entirelyTuning analysis with detection and false-positive rates, above-the-line and below-the-line testing, change approval, periodic re-tuning
Overlays, overrides and qualitative adjustmentsOrganisation’s, entirelyDocumented basis, quantification, approval at the right level, tracking of persistence, sunset criteria
Version upgrades accepted from the vendorShared: the vendor changed the engine, the organisation accepted itRelease notes read, impact assessed, parallel run or test deck, approval to move to production, revalidation where the methodology changed
The vendor’s core algorithm and its training dataVendor’sDevelopmental evidence obtained, limitations recorded, compensating monitoring in place

Auditors should expect to find the boundary drawn in the wrong place in both directions. Organisations exclude configuration from validation because “it is a vendor model”, and they attempt to validate the vendor’s algorithm they cannot see instead of the thresholds they set themselves. The test is a simple one: for each element in the table, who changed it last, with what approval, and what evidence supported the change. The answer for the organisation’s elements is usually a ticket and a spreadsheet, and that is where the findings are.

Compensating monitoring when validation is limited

Where conceptual soundness cannot be assessed, ongoing monitoring and outcomes analysis carry more of the weight, and the framework expects them to be more frequent and more searching for vendor models than for internal ones. Outcomes analysis compares what the model predicted with what happened, on the organisation’s own population, by segment and over time: actual losses against expected, actual fraud against alerts, actual demand against forecast. Population stability monitoring detects the drift that makes a vendor’s calibration wrong for this user: the distribution of inputs and outputs compared with the population the vendor developed on, where that is known, and with the organisation’s own history. Benchmarking runs the challenger alongside. Override and exception analysis watches what users do with the output, since a rising override rate is the earliest sign a model has stopped fitting. And version monitoring tracks what the vendor changed, because a vendor that upgrades the model in a routine release has re-developed the organisation’s model without telling it, and the release notes are the only notice it will get.

Each monitoring strand needs a threshold that triggers action and a route to the model owner and the risk function, in the same way the third-party program’s indicators need an escalation route; the discipline is the same one the management review controls guide applies to any review, which is that a review without a defined trigger for investigation is not a control. For vendor models the thresholds should be tighter than for internal models, precisely because the organisation cannot diagnose the cause of a deviation from inside; it can only notice earlier.

Generative and foundation models: validating the layer you own

Foundation models consumed through an interface are the extreme case of the vendor model: the organisation cannot see the model, cannot obtain its training data, cannot replicate its outputs, and cannot expect the vendor to explain a specific answer. The framework’s response is not to give up but to relocate the model boundary. Everything the organisation builds around the vendor’s model is the organisation’s model, and it is validatable in the ordinary way: the instructions and prompts that shape the outputs, the documents and data the system retrieves and supplies as context, the filters and guardrails that constrain what goes in and what comes out, the human review that sits between the output and the decision, the evaluation set that measures whether the whole assembly does what it is supposed to, and the logging that makes any of it auditable afterwards. The table sets out that layer and the evidence for each element; the EU AI Act’s obligations for high-risk uses, now deferred to December 2027, and the NIST AI Risk Management Framework both point at the same set.

Element of the organisation’s layerWhat it controlsEvidence to validate
Instructions and prompt templatesTask definition, constraints, tone, refusal behaviour, output formatVersion-controlled templates; change approval; the rationale for constraints; test results before and after changes
Retrieval corpus and context dataWhat the model is allowed to draw on; currency and accuracy of the answersCorpus inventory, ownership, refresh cadence, access controls, exclusion of confidential or personal data not intended for the use
Input and output guardrailsBlocked topics, data-leakage filters, toxicity and injection defences, output validation rulesThe rule set, its testing against adversarial cases, monitoring of blocked events
Human review and decision rightsWhich outputs a person must review before they act, and whoThe review points defined by use-case risk; evidence reviews happen; override and correction rates
Evaluation set and acceptance criteriaWhether the assembly performs to standard on the organisation’s own casesA test set built from real use cases with expected outcomes; scores at acceptance and on every change to the model version, prompts or corpus; thresholds for withdrawal
Logging and monitoringAuditability and drift detectionInputs, outputs, versions and reviewer actions logged; periodic sampling for quality; vendor model version changes tracked and re-evaluated

Two rules make this layer defensible. The evaluation set is the validation: without a set of the organisation’s own cases with known good answers, run at acceptance and re-run on every change, there is no evidence that the system works, only that it produces fluent output. And vendor version changes are model changes: when the provider replaces the underlying model, which providers do on their own schedule, the organisation’s assembly has been re-developed, and the evaluation set has to be re-run before the new version reaches users. The organisations that treat foundation models as a utility, like electricity, discover the difference when the utility quietly changes what it says.

Contractual protections that make validation possible

Almost every substitute and monitoring strand above depends on something the vendor has to provide, and vendors provide what the contract requires. The clauses below belong in every contract for a model of the first three types, and the cost of omitting them is discovered at the first validation, when the vendor’s answer to a request for developmental evidence is a price. The third-party program’s contract standard in the lifecycle guide covers the general clauses; these are the model-specific additions.

ClauseWhat it securesTest
Developmental evidence and documentationAccess, under confidentiality, to the model’s design documentation, intended use, development population, limitations and the vendor’s validation results, at onboarding and on each material changeEvidence was requested and received; the file holds it, not a reference to it
Change notification and release notesAdvance notice of methodology, parameter or data changes, with release notes specific enough to assess impact; the organisation’s right to defer an upgradeNotices received in the period reconciled to versions in production; upgrades with impact assessments and approvals
Test-deck and parallel-run rightsThe ability to run the organisation’s own cases through the current and the new version before acceptance, and a test environment to do it inTest decks exist and were run for the last upgrade; results retained
Performance reportingPeriodic reporting by the vendor of the model’s performance on the organisation’s population, where the vendor computes itReports received and reviewed against the organisation’s own outcomes analysis
Audit and regulator accessThe right for the organisation’s auditors and supervisors to examine the vendor’s controls over the model, and to obtain information a supervisor requests about itPresent, exercised or exercisable; supervisors’ requests in the period met
Termination contingencyTransition assistance, data and configuration return, and continued service for a period, so the organisation can replace the model in an orderly wayA contingency plan exists that relies on the clause; the return of configuration has been tested or at least specified
Intellectual property in the configurationThe organisation owns its configuration, assumptions, overlays and the data it supplied, and can extract themClause present; an extract has been produced at least once

Inventory, tiering and the contingency plan for the model’s disappearance

Vendor models belong in the model inventory with the same fields as internal models plus three: the vendor and version, the configuration owner, and the contingency plan. Finding them is the first problem; the models that escape the inventory are the embedded analytics inside business applications, the scores consumed through an interface nobody thinks of as a model, and the end-user tools that call a vendor’s API, which is why the discovery methods in the end-user computing audit guide apply here. Tiering follows the framework’s usual materiality and complexity logic, with one adjustment: a vendor model’s tier should reflect the organisation’s inability to see inside it, so that a black-box score driving a material decision sits at least one tier higher than an equivalent internal model would.

The contingency plan is the item most inventories lack and the guidance has required since 2011: what the organisation would do if the vendor withdrew the model, went out of business, was acquired, or made a change the organisation could not accept. For a configurable engine the plan is usually a migration to an alternative platform carrying the organisation’s own configuration, which is why owning and being able to extract the configuration matters. For a proprietary score it is the challenger model, promoted from benchmark to production, which is a strong reason to build the challenger before it is needed. For an embedded analytic it is the manual or spreadsheet process the analytic replaced, kept documented. The plan needs the same realism as the exit routes in third-party resilience: a named route, a duration, the preconditions in place, and a walkthrough.

The validation and audit program: twelve steps

The steps below serve two purposes: as the validation approach for the model risk function, and as the audit program internal audit runs over that function’s work on vendor models. Sized for one material vendor model, a validation runs 120 to 250 hours; an audit of the vendor-model population across the inventory, sampling three to five models, runs 250 to 400.

1. Confirm the model is in the inventory with vendor, version, configuration owner, tier and contingency plan, and that the inventory’s discovery method would find models of its kind. 2. Obtain the developmental evidence, the vendor’s validation results and the intended-use statement; compare the development population and intended use with the organisation’s actual use and record every difference. 3. Draw the customisation boundary for the model: list every configuration element, its owner, its last change, the approval and the evidence. 4. Validate the organisation’s elements with full rigour: assumptions supported and sensitivity-tested, thresholds tuned with documented detection and false-positive analysis, overlays quantified and approved, data mappings reconciled. 5. Assess conceptual soundness to the extent the evidence allows, and record explicitly what could not be assessed and why. 6. Run the test deck or replication: the organisation’s cases through the model, compared with expectations, a benchmark or the prior version. 7. Perform or review outcomes analysis on the organisation’s population by segment and period; compare with the vendor’s performance reporting. 8. Review population stability and override and exception rates against thresholds; trace breaches to investigation. 9. Reconcile vendor change notices in the period to versions in production; for each upgrade, obtain the impact assessment, the test-deck result and the approval; identify upgrades that changed methodology without revalidation. 10. Read the contract against the seven protections; list what is missing and what has never been exercised. 11. Test the contingency plan: route, duration, preconditions, walkthrough. 12. Conclude with a statement of what is known about the model, what is not, the compensating controls in place for the gap, and the residual risk accepted by the model owner and the risk committee.

Worked example: a bank’s allowance engine and fraud score, validated from the outside

Lakeshore Bancorp, the 9-billion-dollar regional bank used across this site, runs its expected credit loss allowance on a vendor platform and scores card transactions with a vendor fraud model embedded in its card processor’s service. Both had been in the model inventory for years and both had been “validated” in the sense that a validation report existed. Internal audit’s review of the vendor-model population, 320 hours across four models, took the two as its worked cases, and they illustrate the two ends of the visibility spectrum.

The allowance platform was a configurable engine with a documented methodology, and the boundary test showed that the validation had been pointed at the wrong side of it. The prior validation had spent most of its hours on the vendor’s methodology, which was documented, published and used by hundreds of institutions, and a few pages on the bank’s configuration, which was where the bank’s model actually lived: 23 segments, loss-rate assumptions sourced from the bank’s own history with adjustments, a set of qualitative factors, and macro scenarios chosen from the vendor’s library. The year-end control deficiency evaluation had already identified the qualitative factors as a significant deficiency, with a 2.5-million-dollar range of reasonable outcomes and thin support; the vendor-model review traced the cause to the boundary: the factors had been treated as an input to a vendor model rather than as the bank’s own model, and so had escaped the assumption governance applied to everything else. Two vendor version upgrades in three years had been accepted on the vendor’s release notes without a parallel run; one had changed the default treatment of prepayments. The remediation drew the boundary in writing, moved the configuration under the model risk function’s full validation cycle, established a test deck of forty loans for every future upgrade, and added the configuration-ownership and change-notification clauses at the platform’s renewal.

The fraud score was the opposite case: a proprietary model inside the processor’s service, with nothing visible but the inputs, the score and reason codes, and the thresholds the bank set. Nobody had validated it because “it is the processor’s model”, and nobody had validated the thresholds because they were “just settings”. Outcomes analysis on twelve months of the bank’s own transactions, the first ever done, showed the model detecting confirmed fraud well at high scores and the bank’s thresholds set so conservatively that alerts ran at eleven times the confirmed-fraud rate, with the operations team clearing most of them by rule of thumb; the override analysis showed analysts approving transactions at the highest score band at a rate that made the band meaningless. The developmental evidence, requested from the processor for the first time, arrived within a month under confidentiality and showed the model was developed on a large-issuer population with a higher average ticket than the bank’s. The remediation was the compensating-monitoring set: a quarterly outcomes report by score band and segment, a tuning exercise on the thresholds with detection and false-positive rates documented and approved, an override threshold that triggers review, a simple internal challenger scorecard built on the bank’s own history to benchmark the vendor score, and the processor’s change notices added to the model owner’s monitoring calendar. The bank’s residual-risk statement said, accurately, that it could not validate the processor’s algorithm and had chosen to watch its outputs more closely than it watched any model it owned, which is what the guidance means by compensating for the limits of validation.

Related guides

Comments

Leave a Reply

Discover more from internalauditguide.com

Subscribe now to keep reading and get access to the full archive.

Continue reading