SR 26-2 and GenAI Model Risk: What Banks Must Build Now
US regulators retired SR 11-7 and left generative AI outside the new model risk guidance. Here is what changed, what survives, and the governance stack banks and insurers should build now.

On April 17, 2026, the Federal Reserve, OCC and FDIC retired SR 11-7, the fifteen-year-old rulebook that defined model risk management for US banking, and replaced it with a revised interagency framework. The new guidance explicitly excludes generative and agentic AI from its scope, which means the models your institution is deploying fastest are now the ones with the least prescriptive rulebook behind them. The right response is not relief; it is to build a GenAI governance stack deliberately, using the principles that survived the rewrite, before examiners and your own board ask for it.
What actually changed on April 17, 2026
The Federal Reserve issued SR 26-2, Revised Guidance on Model Risk Management, and the OCC issued the matching Bulletin 2026-13. Together they do four concrete things.
First, they supersede SR 11-7 (April 2011) and SR 21-8, the 2021 interagency statement on model risk for BSA/AML systems. The OCC went further and rescinded its Comptroller's Handbook model risk booklet, the 1997 credit scoring guidance, OCC 2011-12 and OCC 2021-19. Fifteen years of accumulated MRM paper collapsed into one document.
Second, the definition of a model tightened. A model is now a complex quantitative method, system or approach that applies statistical, economic or financial theories to turn input data into quantitative estimates. Simple arithmetic, such as spreadsheet calculations, and deterministic rule-based software with no statistical theory behind it are out. A chunk of what sat in bank model inventories for a decade no longer needs to be there.
Third, the cadence changed. The old regime produced a de facto annual revalidation of every model in the inventory. The new framework calibrates oversight to model materiality and to the size and complexity of the institution. The OCC had already signalled this in 2025 when it clarified that community banks are not required to perform annual model validation. Validation effort now follows risk, not the calendar.
Fourth, the guidance is explicitly non-binding. The guidance text itself states it sets no enforceable standards and that non-compliance will not, by itself, draw supervisory criticism. It is expected to be most relevant to organizations over $30 billion in assets, though smaller institutions with heavy model exposure should still read it.
The most important line in the bulletin: GenAI is out of scope
The OCC bulletin states it directly: generative AI and agentic AI models are novel and rapidly evolving, and they are not within the scope of the guidance. Your credit scorecards, stress testing models and AML transaction monitoring systems get a revised rulebook. Your LLM-powered credit memo drafter, claims summarizer, customer service agent and coding assistant get nothing.
This is not deregulation of AI. Three signals make that clear:
- The agencies announced they will issue a request for information on model risk management generally, and on banks' use of AI including generative and agentic AI, in the near future. That RFI is how the next framework gets written, and it will be shaped by what examiners see in the meantime.
- Supervisory authority over unsafe or unsound practices is untouched. The guidance being non-binding does not stop an examiner from criticizing a bank whose GenAI system misadvises customers or leaks data.
- Boards remain accountable. When an AI-driven decision harms a customer, "the regulators had not issued guidance yet" is not a defense that survives a board meeting.
The practical reading, as one detailed analysis of the transition puts it, is that the carve-out is an obligation, not a relief. You now own the design of GenAI governance in full.
What survives from SR 11-7 and still maps to GenAI
SR 11-7 is rescinded, but the principles it crystallized were sound, and most of them transfer to generative systems with some translation work:
- Model inventory. You cannot govern what you have not catalogued. Every GenAI use case, including embedded AI inside vendor SaaS, belongs in a register with an owner, a purpose and a materiality rating.
- Effective challenge. Someone with independence, expertise and standing must be able to say no. The new framework judges challenge on its quality rather than on a rigid org chart, which is an opening to build a leaner, sharper review function.
- Validation becomes evaluation. For a deterministic model, validation asks whether the math is conceptually sound. For an LLM application, the equivalent is a pre-deployment evaluation harness: a fixed test set, scored outputs, and documented pass thresholds.
- Outcomes analysis and ongoing monitoring. SR 11-7 always required checking whether outputs matched reality in production. For GenAI this becomes the primary control, because the model changes under you and the input distribution drifts.
- Vendor and third-party oversight. The revised guidance keeps a full section on validating vendor products. For GenAI, where the model is almost always someone else's, this becomes the dominant discipline.
- Materiality and aggregate risk. Not every use case deserves the same rigor, and many small AI uses can sum to a large exposure. Both ideas carry straight over.
The NIST AI Risk Management Framework and its Generative AI Profile, published in July 2024, are the natural scaffold for the translation. The profile organizes GenAI risk work around the Govern, Map, Measure and Manage functions and highlights four practical considerations: governance, content provenance, pre-deployment testing and incident disclosure. It is voluntary, it is concrete, and it gives your team a shared vocabulary while the banking regulators write their next letter.
A risk-tiered governance stack for generative AI
Here is a build sequence that applies the surviving principles to GenAI without waiting for the RFI.
Step 1: Build the GenAI inventory
List every generative AI use case in production, in pilot and embedded in vendor tools. For each, record the business owner, the model and vendor behind it, the data it touches, the decisions it influences, and the user population. Expect this exercise to find use cases nobody approved.
Step 2: Tier by materiality
Assign each use case a tier based on what happens when it is wrong:
- Tier 1, high materiality. Output feeds credit, claims, pricing, compliance or customer communications directly. A wrong answer creates financial loss, regulatory breach or customer harm.
- Tier 2, moderate. Output is reviewed by a person before it matters, but volume is high enough that review quality can degrade.
- Tier 3, low. Internal productivity uses with no direct customer or financial decision impact.
Step 3: Set controls per tier
- Tier 1: pre-deployment evaluation on a fixed test set with documented thresholds, human review of outputs, logging of every input and output, drift monitoring, a named accountable owner, and a rollback plan.
- Tier 2: pre-deployment evaluation, sampled output review, logging, and quarterly monitoring checks.
- Tier 3: registration in the inventory, acceptable-use policy, and an annual re-check.
Step 4: Stand up pre-deployment testing
For every Tier 1 and Tier 2 use case, build an evaluation set from real historical work product, score the system against it, and write down the pass bar before go-live. This is the GenAI equivalent of validation, and NIST's profile lists pre-deployment testing as one of its four core considerations.
Step 5: Monitor outcomes continuously
Track accuracy against sampled live traffic, override rates by reviewers, complaint and incident signals, and vendor model version changes. When the vendor ships a new model version, your evaluation harness reruns. Monitoring is where the revised guidance puts the weight for frequently updated and vendor models, and GenAI is nothing but frequently updated vendor models.
Step 6: Write the incident and disclosure path
Define what counts as an AI incident, who gets paged, how customers are remediated, and what gets documented. NIST lists incident disclosure as a core consideration; your regulator will treat a well-documented incident response as evidence of serious governance.
The governance checklist
Before any Tier 1 use case goes live, confirm: it is in the inventory with a named owner; an evaluation set exists and the system passed it; human review is in the loop; full logging is on; the vendor contract covers model change notification; an incident path is written; the board risk committee has seen the materiality rating.
A worked example: governing a claims summarizer
An illustrative example shows how the tiers translate into numbers. Consider a mid-sized insurer deploying an LLM summarizer that reads claim files and produces a two-paragraph summary and recommended next step for adjusters. Roughly 500 claims a week flow through it.
Because the output influences claims decisions, this is Tier 1. The governance stack looks like this:
- Evaluation set. 400 historical claims with adjuster-written gold summaries, built once and version-controlled. The system must hit a factual consistency score of at least 98 percent against the gold set, with zero tolerance for invented policy terms, before go-live.
- Human review. Adjusters see the summary as a draft, never as a decision. Their edit rate is itself a monitored metric; if more than 20 percent of summaries need material edits in a week, the system is flagged.
- Sampled QA. 5 percent of live summaries, about 25 a week, get full human re-review against the source file. Findings feed a monthly report to the model owner.
- Vendor change control. The contract requires 30 days notice of model version changes. Each new version reruns the 400-claim evaluation set before it is allowed in production.
- Logging and incident path. Every prompt and completion is retained. A fabricated policy clause is a reportable incident with a defined remediation workflow.
The ongoing cost of this stack is modest next to the exposure: the evaluation harness is a few weeks of engineering, the sampled QA is roughly one adjuster-day a week, and the monitoring is dashboards on data you already log. What it buys is an answer to the examiner question that is certainly coming: show me how you know this system works.
For the underlying build choices, such as whether the summarizer should rely on retrieval over your policy documents or a fine-tuned open-weight model, our decision guide on RAG versus fine-tuning for banks and insurers walks through the trade-offs, and the enterprise AI implementation guide covers the path from pilot to production.
Vendor AI: the evidence to demand
Almost every GenAI system in a bank or insurer runs on a vendor model. The revised guidance kept the vendor oversight sections of the old framework, and for GenAI they become the main event. NIST's profile is specific here: it recommends contracts and SLAs that specify content ownership, usage rights, quality standards, security requirements and content provenance.
Ask every AI vendor for five things, in writing:
- Model change notification. How far in advance, and through what channel, you learn about version changes, deprecations and default behavior shifts.
- Evaluation artifacts. The vendor's own test methodology and results for the capability you are buying, not a marketing benchmark.
- Data handling terms. Whether your inputs train anything, where they are processed, retention periods, and subprocessors.
- Incident and audit rights. Your right to logs, to security documentation, and to notification when the vendor has an incident that touched your data.
- Exit terms. How you get your fine-tunes, embeddings, prompts and evaluation data out if you leave.
A vendor that cannot answer these is telling you something about their own governance maturity.
Common mistakes
- Treating the rescission as permission to relax. The principles survived; only the calendar-driven bureaucracy went. Banks that gut their validation function will rebuild it under exam pressure.
- Leaving GenAI out of the model inventory. Out of scope of the guidance does not mean out of scope of your risk function. Unregistered AI use is the first thing the coming RFI response will force you to admit.
- Validating once at go-live. A point-in-time eval is the start of the control, not the control. Without outcome monitoring you will not notice the vendor model update that quietly degrades your summarizer.
- No evaluation test set at all. "We tried it and it looked good" is not validation. Without a fixed, versioned test set you cannot detect drift or regression.
- Copying the SR 11-7 paperwork onto GenAI. A 60-page model risk pack written for a logistic regression does not describe an LLM agent. Write governance documents that match how these systems actually fail: confabulation, prompt injection, data leakage, silent model updates.
- Ignoring aggregate exposure. Ten low-materiality chatbots sharing one vendor and one data pipeline are a correlated risk, not ten small ones.
What this means in Kenya and East Africa
Kenyan institutions should read the US shift as a preview, not a curiosity. The Central Bank of Kenya's survey of the banking sector, conducted in March 2025 and published in its 2025 Bank Supervision Annual Report, found that half of surveyed institutions have adopted AI, with commercial banks at 66 percent. The leading uses are exactly the high-materiality ones: credit risk assessment at 65 percent of adopters, cybersecurity at 54 percent and customer service at 43 percent.
The governance gap is stark. Only 30 percent of institutions have a formal AI strategy, and among those that have adopted AI, 59 percent have no AI policy at all. Meanwhile 93 percent of respondents asked CBK to issue AI guidance covering governance, risk management and incident reporting.
CBK is moving. On September 10, 2026 it released revised draft prudential and risk-management instruments for public consultation, with comments due November 7, covering model risk, data governance, third-party dependence and digital financial services. Kenyan banks, microfinance banks and digital credit providers that build the inventory, tiering, evaluation and monitoring stack described above now will find the eventual CBK framework familiar rather than disruptive. The same applies to data obligations: our guide to the Kenya Data Protection Act and AI in banking and insurance covers the privacy layer that sits underneath any AI governance program, and we track the broader regulatory picture in our AI governance regulation roundup.
Frequently asked questions
Does SR 26-2 apply to my bank?
It is expected to be most relevant to banking organizations over $30 billion in assets, but the OCC says it can also be relevant to smaller institutions with significant model risk from complex models or non-traditional activities. If you are outside the US, it binds nobody, but it signals where global supervisory thinking is heading.
Is generative AI now unregulated in US banking?
No. GenAI is outside the scope of this particular guidance, but general safety-and-soundness authority, consumer protection law and third-party risk rules all still apply. The agencies have also announced a coming request for information specifically on AI, which is the first step toward dedicated expectations.
What replaces annual model validation?
A risk-based cadence. High-materiality models still get deep, frequent validation; lower-materiality, frequently updated and vendor models lean more on ongoing monitoring and outcomes analysis. The frequency follows the model's materiality, not a fixed annual calendar.
Can we keep using our SR 11-7 documentation templates?
The principles still work, but the artifacts need rewriting for generative systems. A validation report template built for deterministic statistical models does not capture confabulation rates, prompt injection resistance, evaluation set composition or vendor model version control, which are the things that actually go wrong with LLM applications.
What framework should we use for GenAI governance in the meantime?
The NIST AI Risk Management Framework and its Generative AI Profile (NIST AI 600-1) are the most concrete voluntary scaffold available: Govern, Map, Measure and Manage, with specific attention to governance, content provenance, pre-deployment testing and incident disclosure. ISO/IEC 42001 is the certifiable management-system alternative if your board or clients want attestation.
Do embedded AI features in vendor software count?
Yes. If a CRM, core banking platform or productivity suite your staff use ships an AI feature that touches customer data or influences decisions, it belongs in your inventory with a materiality tier. Third-party AI is still your risk under every outsourcing and third-party framework your regulator enforces.
Where to go from here
The institutions that will handle the coming AI supervision wave comfortably are the ones that treat this window, between the old rulebook's retirement and the new one's arrival, as build time. If you want a working version of the stack described here, our enterprise AI practice for banks and insurers designs and deploys exactly this: model inventories, risk-tiered evaluation harnesses, private model deployment, and the governance documentation that stands up to examiners, and our Claude for Financial Services work puts governed AI agents into research, KYC and reconciliation workflows on top of it.
About AI Agents Plus Editorial
AI automation expert and thought leader in business transformation through artificial intelligence.

