Model risk is the possibility of adverse decisions or financial consequences from incorrect, misused, or poorly governed model output.
Model risk is the possibility of adverse financial, operational, customer, compliance, or strategic consequences from decisions based on incorrect, misused, or poorly governed model output. It can arise from flawed design, weak data, coding errors, inappropriate assumptions, use outside the intended scope, inadequate validation, or failure to respond when performance deteriorates.
A model can be mathematically correct and still create model risk if it answers the wrong question, uses unrepresentative data, hides material limitations, or is applied to a different product, population, or market condition than intended.
Definitions vary by organization and regulatory framework. A model generally transforms input data through quantitative methods, assumptions, or rules to produce estimates, classifications, forecasts, valuations, risk measures, or decision support.
Examples include:
Spreadsheets, rules engines, scorecards, and expert-judgment processes may or may not fall within a formal model definition. Excluding a tool from the model inventory does not eliminate its risk. The organization still needs controls proportionate to the consequences of error or misuse.
| Source | What can go wrong | Example evidence |
|---|---|---|
| Conceptual design | Theory, variables, structure, or assumptions do not represent the intended relationship | Development documentation, challenger analysis, sensitivity tests |
| Data | Inputs are inaccurate, incomplete, stale, biased, or unrepresentative | Data lineage, quality tests, missing-value treatment, population comparison |
| Implementation | Code, configuration, interfaces, or reports differ from the approved design | Code review, parallel run, reconciliation, change records |
| Estimation | Parameters are unstable, overfit, or based on an unsuitable sample | Out-of-sample testing, uncertainty ranges, benchmark results |
| Use | Output is applied outside its approved purpose, population, horizon, or conditions | Use inventory, policy, user procedures, decision records |
| Interpretation | Users treat an estimate as precise or ignore limitations and uncertainty | Training, reports, disclosures, override rationale |
| Change | Product, customer, market, regulation, or data-generating process shifts | Drift monitoring, stability analysis, change assessment |
| Aggregation | Individually reasonable outputs interact inconsistently across models | Reconciliation, dependency map, enterprise scenario testing |
| Third party | Vendor logic, data, updates, or limitations are not sufficiently understood | Contract, technical documentation, validation access, contingency plan |
| Risk | Main focus | Relationship |
|---|---|---|
| Model risk | Incorrect or misused model output | Can produce poor decisions even when systems run as designed |
| Operational risk | Failed people, processes, systems, third parties, or external events | Coding, deployment, access, and process failures can create model risk |
| Market risk | Loss from prices, rates, spreads, currencies, or volatility | A market-risk model can understate the underlying market exposure |
| Credit risk | Borrower or counterparty failure | A credit model can misestimate default, exposure, or recovery |
| Data risk | Unfit, unavailable, inaccurate, or poorly governed data | Data weakness is a major source of model risk |
| Decision risk | Poor judgment or governance around a decision | A sound model can be overridden or interpreted poorly |
These categories overlap. Classification is less important than identifying the failure pathway, owner, affected decision, and control response.
Document the decision, users, products, population, geography, horizon, outputs, dependencies, and consequences of error. A simple model used for a major capital or pricing decision may deserve more control than a complex model used only for exploratory analysis.
Explain the theory, assumptions, variables, data, transformations, estimation method, limitations, uncertainty, and intended use. Development evidence should allow a qualified reviewer to reproduce the logic and challenge alternatives.
Confirm that code, data pipelines, interfaces, calculations, reports, permissions, and downstream uses match the approved design. Reconcile test output and preserve version and change evidence.
Validation should be sufficiently independent from development and use. Common elements include:
Validation depth should be proportionate to the model’s materiality, complexity, uncertainty, and use.
Approval should specify permitted uses, limitations, performance thresholds, monitoring, validation frequency, open findings, overlays, and conditions that require restriction or redevelopment.
Monitor input drift, output stability, observed outcomes, overrides, exceptions, data quality, user behavior, and changes in market or business conditions. A model can remain technically unchanged while its environment changes materially.
Material changes should trigger review, testing, documentation, and approval. Models with unresolved severe weaknesses may require use restrictions, conservative adjustments, replacement, or retirement.
Assume a lender developed a default model using seasoned salaried borrowers. The lender later expands rapidly into newly formed small businesses but continues using the same model and cutoffs.
The model may still execute exactly as coded, yet model risk increases because:
A sound response could include restricting use, developing a separate segment, applying conservative overlays, increasing manual review, collecting new outcome data, testing sensitivity, and obtaining validation and approval before broader reliance.
The overlay does not eliminate model risk. It needs a documented basis, owner, amount or method, approval, monitoring, and exit criteria.
Ask:
These methods answer different questions:
| Method | Question | Limitation |
|---|---|---|
| Back-testing | Did predictions align with subsequently observed outcomes? | Outcomes may be delayed, sparse, or influenced by intervention |
| Benchmarking | How does output compare with an alternative method or external reference? | The benchmark can also be weak or differently scoped |
| Sensitivity analysis | Which inputs and assumptions drive the output? | Small changes may not represent structural breaks |
| Stress testing | How does the model or decision behave under adverse assumptions? | Results depend on scenario design and model behavior outside observed data |
| Challenger model | Would another defensible approach materially change conclusions? | Multiple models can share data and conceptual weaknesses |
Passing one test is not proof of overall validity. Evidence should be considered together and in the context of intended use.
Complexity can make development and validation harder, but model risk is not limited to opaque methods. A simple rule can be materially wrong, while a complex model can be controlled if its use, data, behavior, limitations, and outcomes are understood.
For machine-learning and vendor models, reviewers may need to address:
Limited access to vendor intellectual property does not remove accountability for the decision.
The Federal Reserve’s April 2026 guidance superseded SR 11-7 for its stated supervisory scope and identifies the banking organizations to which it is most relevant. These banking materials are useful reference points, not universal requirements for every model, company, investor, or jurisdiction.
This article provides general financial education. It is not personalized banking, investment, credit, accounting, technology, legal, regulatory, statistical, or model-validation advice.