Industry Self-Regulatory Initiatives for Responsible AI in Financial Services
Banks are filling a regulatory vacuum left by Washington with their own AI governance standards.

Two regulatory moves in early 2026 settled a question the banking industry had been asking for years: who writes the rules for AI in finance? The answer, at least from Washington, turned out to be nobody new. In March, the White House released its National Policy Framework on AI, and Principle V told Congress flatly not to build a federal rulemaking body for AI. Deployment questions would run through existing regulators and through standards the industry writes itself. Principle VII went further, urging Congress to preempt state AI laws that create "undue burdens," which reads as a push toward one national floor rather than fifty different ones.
Then, on April 17, the Fed, the OCC, and the FDIC issued SR 26-2, guidance that is explicitly non-binding and sets no enforceable standard. Buried in a footnote (footnote 3, to be exact) is maybe the most consequential line in the whole document: generative and agentic AI are carved out of scope entirely, described as "novel and rapidly evolving." So the newest, fastest-growing category of AI in banking just got told, formally, that the formal rulebook doesn't apply to it yet.
That leaves a gap. And gaps in banking don't stay empty. They get filled by whoever's willing to write something down first.
Standards bodies and what each demands
Six efforts are worth knowing by name, because each one is shaping what examiners ask about and what boards get asked to sign off on, even without a binding mandate behind it.
The Financial Stability Board put out a consultation report called "Sound Practices for Responsible Adoption of Artificial Intelligence." It lays out 12 sound practices, backed by case studies from actual institutions, covering AI governance across the full lifecycle, from procurement through retirement. It's aimed squarely at boards and senior management. The FSB says outright that it wants this used as a reference point when firms make decisions about strategy, technology adoption, and risk. It's still a consultation, technically not binding on anyone. But this is the body that coordinates financial stability policy globally, so when examiners walk into a room, they tend to know what's in it.
SR 26-2, paired with OCC Bulletin 2026-13, is the biggest rewrite of model risk management guidance in over a decade. It replaces SR 11-7 and SR 21-8, the two documents banks have built entire validation functions around for a long stretch of years. The new guidance introduces a materiality threshold aimed at banks holding $30 billion or more in assets. But the carve-out for generative and agentic AI is conspicuous. Formally, it creates no enforceable requirement for these systems.
It gets interesting, though: examiners aren't behaving like the carve-out exists. In practice, they're asking about kill-switch capability, data access boundaries, and vendor oversight for agentic systems in exams right now, regardless of what the footnote says. SR 26-2 draws a line around where the formal model risk framework stops. Everything past that line is something each institution has to build governance for on its own, without a template to copy.
The CRI Financial Services AI Risk Management Framework, released by the Treasury Department, tries to close part of that gap. It was built jointly by the Cyber Risk Institute and the Financial Services Sector Coordinating Council, under Treasury's AI Executive Oversight Group. Functionally, it's the NIST AI RMF translated into financial-services language and mapped onto the CRI structure banks already know. It comes with practical tools including an AI Adoption Stage Questionnaire, so an institution can locate itself honestly on the adoption curve, and a Risk and Control Matrix with 230 distinct control objectives. Treasury also released a shared AI Lexicon alongside it, which sounds like a small thing until you realize how much friction gets created when a regulator and a bank use the word "explainability" to mean two different things. It's voluntary. But it scales down to small credit unions and up to global banks, which is more than most frameworks manage.
Agentic AI's place in banking practice and the governance lag behind it
Adoption is not waiting for anyone. The Cambridge Centre for Alternative Finance found 81% of financial services firms adopting AI at some level, with 40% already at the advanced "Scaling" or "Transforming" stages. Regulators, by comparison, are at around 20% at those same advanced stages. Regulators, by comparison, are at around 20% at those same advanced stages, a 20-point gap between industry and oversight. That's an industry that's lapped its overseers.
Cornerstone Advisors' What's Going On in Banking 2026 report found 59% of credit unions have already deployed generative AI in some form. Credit unions, historically the slower-moving cousin of the banking world, are not sitting this one out.
Agentic AI specifically, meaning systems that don't just generate text but take actions and chain decisions together, is further along than most people outside the industry probably assume. The same Cambridge Centre report puts 52% of financial services firms as already rolling out agentic AI in some capacity. More than half. The guidance excludes a category of technology that a majority of the industry has already started deploying.
Platform activity backs this up with specifics. Fiserv launched agentOS, built on AWS Bedrock AgentCore with OpenAI as a strategic collaborator, and Early institutional partners are already piloting it, with additional institutions co-developing what comes next. FIS unveiled a Financial Crimes AI Agent, with partner institutions first in line and broader availability planned. Jack Henry has been building out AI capabilities for its roughly 7,400 community bank and credit union clients, and early adopters are reporting time savings of up to 70% on routine administrative work. Eltropy launched an agentic AI platform built specifically for credit unions. Backbase followed with an AI-native Banking OS designed as an operating layer above existing infrastructure.
So the technology is live, in production, moving money and making decisions inside real institutions. The governance meant to oversee it is still marked "novel and rapidly evolving" in a regulatory footnote. That mismatch is the story here, and it's the reason self-regulatory frameworks matter more this year than they have in a long time.
What the frameworks require institutions to build internally
Stripping away the branding differences between the FSB's sound practices, SR 26-2, and the CRI framework reveals four requirements common to all of them.
Accountability. Someone at the executive level has to own every AI outcome, not a committee, not a vague "the model team." SR 26-2, the FSB's sound practices, and a UK individual accountability and certification regime all converge on this point from different directions.
Transparency. An institution has to be able to explain how a specific model reached a specific decision. The FCA's David Geale said explainability and governance for AI models remain "non-negotiable," even in a landscape where no new rules have arrived. That's a regulator saying the absence of a binding rule doesn't mean the absence of an expectation.
Auditability. Every automated action needs to leave behind a record regulators can actually use, not a summary log but something closer to an immutable trail. This is where MAS's SAFR framework operationalizes the idea down to the level of individual actions, and it's also what SR 26-2 examiners are asking for in practice, scope carve-out or not.
Continuous validation. Testing a model once at launch and calling it done doesn't hold up anymore. SR 26-2 pushes lifecycle thinking: ongoing testing rather than a point-in-time check. UK firms have raised this exact concern with their regulators: traditional model risk management approaches, built for models that don't change their own behavior, don't scale to generative and agentic AI.
Putting those four pillars together starts to form a minimum viable governance scaffold, drawn from what SR 26-2 examiners are actually asking about and what the broader industry guidance recommends:
- A complete inventory of every GenAI tool and agent running in production. Research has found that only a small share of organizations have full visibility into their own AI footprint, which means the majority don't know everything that's running. Governance is impossible without that inventory first.
- Monitoring built around how these systems actually break: hallucination, adversarial prompting, and evaluation approaches like LLM-as-a-judge, rather than the failure modes of a traditional statistical model.
- Full source traceability on every automated output, so an examiner (or a bank's own risk team) can trace a decision back to what produced it.
- A human in the loop at every material decision point, with clear lines for who's accountable when that human signs off.
- Documentation of actions, recommendations, responses, and exceptions that can't quietly be edited after the fact.
- Independence between the system doing the work and the system validating that work. The same model shouldn't grade its own homework.
- Change documentation kept audit-ready at all times, not assembled the week before an exam.
The CRI framework's 230 control objectives give this scaffold a practical structure to sit inside, and the Adoption Stage Questionnaire helps an institution figure out which controls actually apply given where it sits on the deployment curve. A community bank piloting a single customer-service agent doesn't need the same control depth as a much larger institution running agentic workflows across payments and fraud detection, and the questionnaire is built to reflect that difference rather than pretend every institution needs the same thing.
MAS's SAFR framework pushes one layer deeper for institutions actually running agentic systems: it specifies how an agent's actions get authorized before they happen, what conditions trigger human review at the moment a consequential decision is about to fire, and what has to be captured in the record of that decision. For any institution running agentic payments or transfer automation, this is the part of the landscape that deserves the closest read, because it's the one framework asking questions at the level where the money actually moves.
Accountability structure differences between traditional model risk management and agentic AI governance
SR 26-2 and its predecessors, SR 11-7 and SR 21-8, were built for a specific kind of model: discrete, human-designed, running fixed decision logic that doesn't change once it's validated. An agent that adapts its behavior mid-workflow and chains several decisions together in sequence was never the thing that framework had in mind.
That mismatch creates a problem that doesn't have a clean name yet, but appears constantly in practice: compounding risk. An individual agent skill might be entirely safe in isolation, tested and approved on its own terms. Combining four or five of those skills into a multi-step workflow raises new risks that none of the individual reviews would have caught. SR 26-2 doesn't address this. Neither does the FSB's sound practices document, nor the CRI framework. Of everything covered here, only SAFR engages with this problem at the runtime level, in the moment an agent is actually stringing decisions together.
UK firms said during those February 2026 central bank roundtables that traditional model risk management and validation approaches can't scale to widespread generative and agentic AI deployment. Validation built for one model at a time doesn't hold up against an agent orchestrating several models across a live workflow. The old framework was solving a different problem. It's just an acknowledgment that the framework was solving a different problem.
Speed makes this worse. Agentic systems can cascade into failure before a human reviewer has any real chance to step in and stop it. Which means the accountability model has to change. Traditional model risk management is largely retrospective: validate the model, deploy it, audit after the fact. Agentic AI needs something prospective, authorization and boundary-setting decided before the action happens, not just a record checked after. That's the entire logic behind SAFR's runtime authorization model, and it's a meaningfully different posture from anything SR 26-2 asks for today.
This is also where the industry's own tools are starting to reflect the shift. Platforms built to execute real banking transactions, through voice or text, while keeping institutional controls and full audit trails intact, are effectively built around the same questions SAFR is asking: who authorized this action, what triggered human review, and what does the record show. PaymanAI is one example of a platform operating in that space, built around executing transactions under the kind of runtime authorization and audit logic that agentic governance frameworks are starting to formalize.
None of the frameworks covered here are mandatory. That's precisely why they matter. When the binding rule hasn't caught up to the technology, the standards an industry writes for itself become the actual floor examiners measure against, whether or not any statute says so.



