OCC Model Risk Guidance Applied to Agentic AI Systems
Regulators exclude agentic AI from model risk rules, leaving banks to build their own controls.

In April 2026, the OCC, the Federal Reserve, and the FDIC replaced SR 11-7 with SR 26-2, and the new guidance does something SR 11-7 never had to do: it draws a line around agentic AI and says, explicitly, that this falls outside formal model risk scope. It functions as an assignment: banks now have to build their own controls for the fastest-moving part of their AI deployment, without a prescribed path and without the safety net of examiner-approved standards to lean on.
Why the SR 11-7 definition of "model" never fit agentic systems in the first place
SR 11-7, the 2011 guidance that shaped model risk management for a decade and a half, was built around a specific picture of what a model is. A model gets calibrated, gets deployed, and then sits still until the next review. Someone validates it before it goes live. Someone revalidates it on a schedule, usually annually. In between, the thing doesn't change. You can pull a snapshot of its parameters at any point and reconstruct exactly why it produced the output it did.
That picture assumes stable parameters, a known relationship between inputs and outputs, and a human checking the work at each stage before anything happens downstream.
Agentic systems break every part of that assumption. They learn and adapt while running in production, not just between formal review cycles. They don't produce an estimate for a human to act on; they act. And they often act across a chain of steps, where an early decision shapes what happens three or four steps later, so the risk isn't in any single output but in how the decisions compound along the way.
Here's the part that really breaks the old framework: you can't always reconstruct an agent's decision path from a static snapshot, because the system that made the decision may not be the same system it was an hour before. SR 11-7's entire validation rhythm, validate before deployment, revalidate on a calendar, assumes a model that holds still between checkups. An agent that adjusts itself in real time stresses that cadence past the point of usefulness.
This isn't some narrow edge case that only shows up in unusual deployments. The gap between what SR 11-7 was built to govern and what agentic systems actually do is wide enough that bolting SR 11-7-style controls onto an agent gives you a governance framework that looks thorough in a binder and misses the actual risk sitting in front of you.
How far agentic AI has already penetrated banking operations
Here's what makes the timing of this carve-out uncomfortable: agentic AI isn't a future consideration for banks. It's already running. Recent industry surveys indicate that a large majority of banking leaders say their firms use agentic AI in some form, whether through live deployments or active pilots.
Some of this is already at production scale. PaymanAI, for instance, deploys AI agents that execute real banking transactions on a bank's existing rails, with full audit trails built in. Scotiabank's AIDox started in 2018 as a document analysis tool and now autonomously processes client emails in commercial banking, routing them, triaging them, and opening cases without a person kicking off the workflow. Salient runs end-to-end loan servicing for major auto lenders, reportedly with zero customer churn and full conversion from pilot to contract. Visa's Trusted Agent Protocol lets AI search, compare, and pay on a consumer's behalf inside agent-driven checkout flows. BBVA has embedded its banking app directly into ChatGPT for customers in Germany and Italy.
Then there's KYC and AML, arguably the highest-stakes corner of this whole conversation. Banks devote a meaningful share of full-time staff to KYC and AML work alone, yet Interpol estimates the financial industry catches only about 2% of global financial crime flows, despite years of rising compliance spend. That gap is exactly why agentic AI looks so attractive here: the productivity case writes itself.
McKinsey's 2025 baseline puts numbers on the stakes. In the most likely scenario, agentic AI adoption could cut costs by 15 to 20%. Fail to adapt, and the roughly $1.2 trillion global banking profit pool could shrink by as much as 10% over the next five to ten years.
This is a production-scale risk, governed right now by frameworks that regulators have said, in writing, are incomplete.
The governance gap that banks must now close on their own
Put plainly: SR 26-2 tells banks to govern agentic AI, but doesn't tell them how. The old toolkit was built for a different animal.
Recent surveys of financial institutions suggest that roughly a third have already pushed AI or machine learning into production. Only a small slice of those describe their AI strategy as well-defined and properly resourced — a documented lag between what's deployed and what's governed. That's a documented lag between what's deployed and what's actually governed.
Credit unions face a stranger version of this. The NCUA doesn't have a model risk rule to carve anything out of in the first place, so examiners reach AI through whatever supervisory lens already exists. Policymakers have recommended that NCUA build out model risk guidance, and NCUA staff concluded that doing so would require formal rulemaking that hasn't happened. For now, the NIST AI Risk Management Framework is the closest thing to an operative reference.
What does agentic governance actually need to answer that the old model risk playbook never had to touch? A few questions worth sitting with: What is the agent allowed to do, and how are those limits enforced while it's running, not just written down somewhere as policy? Who's accountable when a multi-step autonomous action produces a result nobody wanted? How fast can the agent be reined in or shut off once its behavior drifts? And how do you catch an error before it spreads across a workflow the agent is running at scale?
Someone might argue banks should just wait for clearer rules before answering any of this. That raises an obvious problem, though: institutions are being asked to judge the adequacy of their own governance for systems regulators admit they haven't fully mapped out yet. The forthcoming request for information will eventually turn into guidance, but banks running agents today don't get to pause until it does.
Where payments and transfer automation concentrate the compliance exposure
Payments are where this gets sharpest. Older automation followed rules: if X happens, trigger Y. Agentic AI interprets a goal, makes a judgment call partway through a task, and can initiate a payment without a person walking it through each step.
The timing compounds this. Fedwire adopted ISO 20022 in July 2025, and Swift ended its coexistence period with the older format that November — transitions that have been widely documented across the industry. Structured, richer payment data is now the baseline for the whole system. That's good news for fraud detection and automation generally, but it also means agents are now acting on more consequential data than before, with more room to get something subtly wrong at scale.
There's also an authorization trail problem worth sitting with. Some emerging approaches use cryptographically signed mandates that capture a user's upfront, scoped instructions, essentially trying to build a paper trail of intent before the agent acts. It's a sensible idea, yet it hasn't been tested in court, and it hasn't been tested before regulators either. Banks adopting this approach are, in a real sense, ahead of settled law.
Zoom out further and policymakers have flagged something even bigger: correlated agent behavior in payment initiation. If a bunch of agents across different institutions start behaving similarly, the way algorithmic trading strategies sometimes herd together in markets, payment flows can synchronize in ways that spike intraday liquidity demand and strain settlement capacity. That risk extends well beyond any single bank's ability to manage it alone.
Emerging supervisory expectations lay out what policymakers will likely require institutions to do in the meantime: keep AI activity logs and audit trails, run real-time monitoring for anomalies in agent behavior and transaction flows, and generate automated alerts when something looks off. That's a fair preview of what examiners will expect even before any formal rule locks it in. An institution that can already show configurable transaction limits, real-time enforcement of those limits, and a complete log of every agent-initiated payment is simply ahead of peers still treating this as a plain automation upgrade.
What audit-ready controls for agentic AI actually require
Here's the most common audit finding showing up in 2026: an AI system generates a credit adverse action notice, and the audit trail only shows the output score, not the inputs and not the reasoning behind it. Because access ran through a shared service account or an API key instead of a specific person, there's no individual to point to either.
Examiners are looking for a few things that a lot of existing logging setups simply don't capture. Every agentic action tied to a real human identity, not a generic service account. Policy enforcement documented at each step of a multi-step workflow, not just checked once at the start. A clear record of what data the agent touched, in what order, and under whose authority it was touching it.
Retention rules add another wrinkle, especially for banks operating across borders. The EU AI Act's Article 12 sets a floor of at least six months for high-risk AI systems. Basel III practice, meanwhile, tends to align retention with the model's full lifecycle, often seven years or more. A bank operating under both has to satisfy the stricter one, every time.
Voice AI is its own flashpoint. The OCC's Fall 2025 Semiannual Risk Perspective, published by occ.gov, calls out AI-related compliance risk directly, and examiners are now asking pointed questions: what exactly did the voice AI say during a collections call, can the bank defend that disclosure, and which conversations did the AI actually handle versus a person? Institutions that ran voice pilots back in 2024 without building a logging framework around them are the ones feeling that scrutiny land now.
There's a consumer protection floor underneath all of this too: clear channels for AI-related complaints, and documented human oversight for consumer-facing decisions, not just an assumption that a person was watching.
The bigger point here is architectural. Audit readiness for agentic AI has to be built into the system itself, so that compliance can move at the same speed and the same scale as the agent does, rather than arriving as paperwork bolted on after the system ships. Institutions building agents on existing banking infrastructure, with configurable controls, full action logs, and SOC 2-certified systems underneath, are already constructing the kind of control architecture examiners will eventually write into formal rules. They're just not waiting for the rule to show up first.
How to structure an internal governance framework before formal guidance arrives
SR 26-2 is principles-based, not prescriptive, so the framework banks build around it should be too. Three layers cover most of the ground.
An authorization layer defines exactly what an agent can do, under what limits, and under what conditions, and enforces those limits while the agent is running, not just in a policy document nobody checks in real time. A monitoring layer watches the agent's behavior continuously in production, with automated alerts when something drifts and a clear, documented path for escalating it. And an audit layer keeps complete, attributable, human-readable records of every action the agent takes, retained as long as the strictest applicable rule requires.
Proportionality matters here too, and SR 26-2 says as much itself: governance intensity should track risk exposure. A low-volume internal workflow agent handling routine document sorting carries far less risk than a customer-facing agent that can initiate payments, and its governance overhead should scale accordingly.
A few things worth documenting now, because the eventual RFI will likely ask for exactly this: how the institution decided which agentic systems sit outside SR 26-2's scope, what alternative governance practices got applied and why they were judged sufficient, and how the framework will flex once formal agentic AI guidance actually lands.
The NIST AI Risk Management Framework is a useful scaffold in the meantime. Both NCUA and the broader interagency posture point toward it, and mapping existing controls onto its Govern, Map, Measure, and Manage structure gives an institution a defensible position even before sector-specific rules exist.
There's a competitive angle worth naming plainly: institutions that build real governance now, before the RFI closes, will walk into an exam with evidence of maturity already in hand. Institutions that wait will be explaining gaps instead.
SR 26-2's carve-out is an assignment to govern well, on the institution's own judgment, without a map handed to them. The banks that treat it that way, rather than as a pass to move slowly, are the ones that will already be standing on solid ground when the formal guidance finally shows up.


