Most community banks believe model risk management does not apply to them. The belief is understandable — the supervisory guidance was written with large institutions in view, and the word "model" suggests something more elaborate than what a $600 million bank operates.
It is also wrong in a specific and consequential way. A bank that calculates its allowance in a spreadsheet, prices loans with a tool the CFO built, screens transactions with a vendor's monitoring system, and measures interest rate risk in a purchased ALM package is running four models. The question is not whether the institution has models. It is whether anyone has established that they produce correct answers.
This post covers the general framework across all models. Our companion post on CECL implementation covers allowance model governance and validation specifically, which is the one most institutions encounter first.
The supervisory definition is broader than intuition suggests: a quantitative method, system, or approach that applies statistical, economic, financial, or mathematical theory and assumptions to process input data into quantitative estimates.
Three components, all of which have to be present: an input component that handles data and assumptions, a processing component that transforms them, and a reporting component that translates results into usable information.
Applied to a community bank, that typically captures:
What is generally not a model: a report that presents data without transforming it, and a straightforward mechanical calculation with no estimation or theory behind it. The boundary is judgment, and the useful discipline is to document the judgment rather than to resolve it by assertion. A tool the institution decided was not a model, with a written reason, is a defensible position; a tool nobody classified is not.
Every model risk program starts with an inventory, and every institution building one for the first time finds models it did not know it had.
The inventory should record, for each model: its name and purpose, its owner in the business, who developed it, whether it is vendor or internal, what it feeds, its date of last validation and the next scheduled date, its risk tier, and any known limitations or compensating controls.
Two features distinguish a useful inventory from a list. Interdependency mapping — which models feed other models — because an error in a deposit decay assumption propagates into the ALM output, the liquidity projection, and the capital plan. And an owner who is a person, not a department, because unowned models are the ones that go stale.
Finding the models requires looking where they hide. The reliable method is to ask, for each material number that reaches the board or a regulatory report, how it was produced and what tool produced it. That question finds spreadsheets that nobody thought of as models, which is where much of the actual risk sits.
At a community bank, the highest-risk model is frequently a spreadsheet: built by one person, undocumented, unversioned, unprotected, and feeding a number that appears in the Call Report.
The characteristic failure modes are well known and rarely controlled. A formula that does not extend to the last row when a row is added. Hardcoded values inside formulas. A copy-paste error in one cell among thousands. Manual inputs with no reasonableness check. No version control, so nobody can establish which version produced last quarter's figure. And key person dependency — one person understands it, and the institution discovers the dependency when they leave.
Proportionate controls exist and are not expensive: lock cells that should not be edited, separate inputs from calculations from outputs, document the logic in the file itself, maintain version history, require a documented independent recalculation of the material outputs at a stated frequency, and have a second person capable of operating it.
The most common misconception in this area, and vendors do not go out of their way to correct it: the vendor's own testing satisfies the institution's validation obligation.
It does not. Supervisory expectations place the responsibility on the institution using the model, and validation of a vendor model has to establish that the model works for this institution's portfolio, products, and market — not that the vendor's development process was sound in general. A monitoring system calibrated on a national customer base can generate an alert volume that is useless at a bank whose customer profile differs materially.
What the institution should obtain and actually read: documentation of the model's development, theory, and intended use; the vendor's own validation or testing results; a clear statement of the model's limitations and the population it was developed on; and information sufficient to understand the logic. Where the vendor treats the methodology as proprietary, the institution's leverage is at contract negotiation and the compensating control is heavier outcomes testing — if the inner workings cannot be examined, the results have to be examined harder.
What the institution must do itself in every case: validate the input data, since the model's output is only as good as what the institution feeds it; assess and document the parameter and threshold settings, which are institution-specific choices and the most common source of poor performance; and perform outcomes analysis to establish that the results are appropriate for this portfolio.
The transaction monitoring case is worth stating plainly because it recurs in examinations. Thresholds and rule settings inherited from a vendor default, never tuned to the institution's own customer base and never revisited, produce either an unmanageable alert volume or a monitoring system that detects very little. Both are findings, and both are validation failures rather than vendor failures.
Validation is not one activity, and the word gets used loosely. Supervisory guidance frames it as three components, and a validation report that addresses only one is incomplete.
Conceptual soundness. Is the model's design and theory appropriate for its purpose? This reviews the methodology, the assumptions, the developmental evidence, the data used to build it, and the documented limitations. It answers whether the approach could produce a correct answer in principle.
Ongoing monitoring. Is the model still working, and is it being used as intended? This covers process verification — that the model runs correctly on current data — benchmarking against alternatives or peer information, and sensitivity analysis of the assumptions that drive the result. This is continuous rather than periodic, and it is the component institutions most often skip entirely.
Outcomes analysis. Do the model's outputs match what actually happened? Back-testing where the model produces predictions with observable outcomes, comparison of estimates to realized results, and analysis of the errors. For an allowance model this means comparing estimated losses to actual charge-offs over time; for a credit scoring model, comparing predicted default rates to realized ones by score band.
The independence requirement runs through all three: validation must be performed by someone who was not responsible for developing or operating the model, with the standing to challenge it and a reporting line that does not run through the model's owner. At a small institution this frequently means a qualified third party, and it can also mean a competent person in a different function — but it cannot mean the person who built the tool.
The agencies issued a statement addressing the application of model risk management guidance at community banking organizations, and its content matters: the framework's rigor should be proportionate to the institution's size, complexity, and the materiality of its models. A community bank is not expected to run a large institution's program.
What proportionality permits: risk-tiering models so that high-risk ones receive full independent validation and low-risk ones receive review; using qualified third parties rather than an internal validation function; validating high-risk models less frequently than annually where change and performance justify it; and documenting proportionately rather than exhaustively.
What proportionality does not permit: no inventory, no owners, no validation of the models that produce material numbers, no documented rationale for the tiering, and no board awareness of model risk. Scaling the program is expected; omitting it is not.
Risk-tiering should reflect the materiality of the model's output, the complexity of its methodology, the consequence of it being wrong, whether it feeds regulatory reporting, and how much reliance is placed on it in decisions. A pricing tool that informs judgment sits lower than an allowance model that produces a Call Report line.
The board approves the model risk policy and should receive periodic reporting on the model inventory, validation status, and any material unresolved findings. It does not need to understand model methodology; it needs to know whether the numbers it relies on have been established as reliable.
Senior management owns the framework, assigns model ownership, and resolves validation findings.
Model owners in the business are accountable for their model's performance, data quality, appropriate use, and documentation.
Validators are independent, and their findings are tracked to closure like any other issue.
Internal audit assesses whether the framework itself operates as designed — not whether the models are correct, which is the validator's role.
Two governance items that materially improve a program. Change control: a model change, a parameter or threshold change, or a new data source should trigger an assessment of whether re-validation is required, and version history should record what changed and when. And use documentation: a model validated for one purpose and then used for another is a common and unmanaged failure — the allowance model repurposed for stress testing, the pricing tool used for portfolio valuation.
Structured coverage is available through the Certificate in Risk Management, the Certificate in Financial and Credit Risk Management, Introduction to Credit Risk Management, and the Certificate in Operational Risk Management.
No inventory, so the institution cannot state how many models it has.
Spreadsheets excluded from scope, which omits the highest-risk models at most community banks.
Vendor models treated as validated by the vendor, with no institution-specific outcomes analysis and no documented assessment of the threshold settings.
Validation performed by the model's owner, which is review rather than validation regardless of how thorough it is.
Only conceptual soundness addressed, with no ongoing monitoring and no outcomes analysis — the two components that would reveal a model that has stopped working.
Validation findings not tracked, so the report identifies limitations that persist unaddressed through the next validation.
Assumptions never revisited after the environment changed, which is the interest rate risk lesson of the last cycle applied to models generally: deposit behavior assumptions calibrated in one rate environment were carried unchanged into another.
No change control, so nobody can establish which version of a model produced a reported figure.
The unifying failure is the one that recurs across every discipline in this series: the institution relies on a number without having established that the number is right, and the reliance is invisible until the number matters.
Models that make or substantially influence decisions about individuals raise a second set of issues beyond accuracy, and they belong in the model risk framework rather than alongside it.
Where a model influences credit decisions, pricing, or account actions, the questions extend to fair lending — whether the model or its variables produce a disparate impact, whether apparently neutral inputs proxy for protected characteristics — and to adverse action notice obligations, which require the institution to state specific principal reasons for a denial. A model whose logic cannot be explained well enough to generate accurate reasons creates a compliance problem independent of whether its predictions are good.
Two disciplines follow. Test for disparate impact as part of validation, not as a separate compliance exercise conducted later. And establish explainability before deployment, because a model that cannot produce a defensible reason for an individual decision is not deployable in a credit context however accurate it is in aggregate.
The general point holds for any model that uses machine learning or is supplied as an opaque service: the harder it is to examine the mechanism, the more weight has to fall on outcomes testing, monitoring for drift, and documented human review of the decisions the model drives.
Yes, proportionately. The agencies issued a statement addressing application of the guidance at community banking organizations, confirming that rigor should scale with size, complexity, and model materiality. Scaling the framework is expected; having no inventory, no owners, and no validation of the models producing material numbers is not.
If it applies assumptions and calculation to input data to produce an estimate, yes — and at most community banks the spreadsheets are the highest-risk models, because they are typically undocumented, unversioned, unprotected, and dependent on one person. Locking cells, separating inputs from calculations, documenting the logic, and requiring an independent recalculation are proportionate controls.
No. The responsibility sits with the institution using the model, and validation must establish that the model works for this institution's portfolio and market. The bank must validate its input data, assess and document its own parameter and threshold settings, and perform outcomes analysis regardless of what testing the vendor performed.
Conceptual soundness — whether the design and theory suit the purpose; ongoing monitoring — whether the model still works and is used as intended, including process verification, benchmarking, and sensitivity analysis; and outcomes analysis — whether results match what actually happened. A validation report addressing only the first is incomplete.
Someone independent of the model's development and operation, with the standing to challenge it and a reporting line outside the model owner's. At a small institution that often means a qualified third party, and it can mean a competent person in another function — but not the person who built or runs the model.
Because rules-based monitoring systems are models, and their thresholds are institution-specific settings that are frequently left at vendor defaults, never tuned to the bank's own customer base, and never revisited. That produces either an unmanageable alert volume or a system that detects very little, and both are validation failures rather than vendor failures.


