NIST AI RMF Applied to Agentic Banking Deployments
How banks can use NIST's framework to govern AI agents executing real transactions.

Agentic AI in banking means an AI system that doesn't just tell you something, it does something: initiates a wire, moves money between accounts, closes out a collections case. That's a different animal from the recommendation engines and fraud scores banks have run for years, where a human still reads the output and decides. Once an AI agent can execute a real transaction with little to no human review at each step, the question stops being "is the model accurate" and becomes "who authorized this, under what limits, and can we prove it after the fact."
The market has already made its bet. The global agentic AI for banking market sat at $4.8 billion in 2025 and is projected to hit $78.6 billion by 2034, growing at a 38.2% compound annual rate. That represents an industry moving fast toward autonomous execution. And yet, per Evident's AI Banking Index for Q1 2026, only 9 of the 50 financial institutions it tracks had actually gotten an agentic AI application into live production or even a real pilot. Everyone's talking about it. Almost no one has shipped it.
So what's the holdup? The models can already plan multi-step tasks, call external tools, and adjust their own approach mid-task, so the technology itself is largely there. What's missing is the governance scaffolding to let a bank trust an agent with real money and defend that trust to an examiner. Traditional model risk management, the kind built on the old SR 11-7 lineage, assumes a model spits out a prediction and a human decides what to do with it. Agentic systems break that assumption at the root: they perceive, decide, and act, often in the same breath. The rest of this piece walks through where regulators currently stand, what NIST's AI Risk Management Framework offers, where it falls short for agents specifically, and how the four core functions of that framework, Govern, Map, Measure, and Manage, apply directly to a bank trying to deploy agents responsibly.
Where the regulatory floor currently sits — and the gap SR 26-2 leaves open
SR 26-2, issued jointly by the Federal Reserve, the OCC, and the FDIC in April 2026, is now the interagency standard for model risk management, replacing the long-standing SR 11-7. Here's the detail worth sitting with: SR 26-2 names generative and agentic AI specifically, and it does so to place them outside its scope. The newest piece of model-risk guidance we have explicitly carves out the exact systems that most urgently need governing.
That carve-out doesn't mean no oversight applies, though SR 26-2 is direct on that point. What it does mean is that the burden shifts onto the institution to figure out which other frameworks fill the gap and to build controls accordingly. Nobody's handing banks a checklist here.
There's a third-party wrinkle too, and it matters a lot given how most banks actually deploy AI. SR 26-2 holds banking organizations fully responsible for validating vendor models, even when the vendor won't hand over proprietary code, training data, or methodology. Since the overwhelming majority of agentic deployments in banking run on third-party platforms, this provision isn't a footnote; it's close to the whole ballgame.
On the sector-specific side, the Financial Services AI Risk Management Framework, built by the Cyber Risk Institute with support from the U.S. Treasury, takes NIST's AI RMF 1.0 control objectives and maps them onto banking-specific risks: bias, model opacity, cybersecurity exposure, operational resilience. It's the clearest translation we have right now of general AI risk language into bank-speak.
Put it together and the message from regulators is consistent, if incomplete: agentic AI needs governance, but nobody's written the detailed rulebook yet. That's both a real obligation and a real opening. Institutions that build a defensible framework now, before the prescriptive guidance lands, are going to be in a much better spot than the ones waiting for someone to tell them exactly what to do.
What the NIST AI RMF is and why financial institutions should care about its 2024–2026 evolution
NIST released the AI Risk Management Framework, version 1.0, in early 2023. Since then it's grown well past its original scope, picking up companion playbooks, sector profiles, and evaluation tools until it's become one of the most widely referenced voluntary AI governance frameworks anywhere in the world.
The framework runs on four core functions, and they're meant to work as a continuous loop rather than a box you check once a year:
- GOVERN: set organizational policy, assign roles, build accountability structures
- MAP: understand risk in the context of a specific system and specific use case
- MEASURE: assess and monitor risk with both quantitative and qualitative methods
- MANAGE: respond to risk, track whether the response worked, adjust
The framing shift from the old SR 11-7 world is worth calling out directly. NIST treats AI as what it calls a socio-technical system, meaning the risk lives across people, use context, and the full lifecycle of the system, not just in whether the model's output is statistically sound. That's a broader lens than banking's model risk teams are used to, and it's exactly the lens agentic systems need.
The framework has kept moving, and each update lands closer to home for banking:
In July 2024, NIST published NIST-AI-600-1, the Generative AI Profile, adding 12 risks specific to generative systems: hallucination, data poisoning, prompt injection, over-reliance, among others. In March 2025, an updated NIST AI 100-2 named AI agents as a distinct threat surface for the first time. By December 2025, NIST's Cyber AI Profile tailored the Cybersecurity Framework 2.0 to AI-specific security risks; American Banker reported that banks received new federal guidance on AI cyber risk around that same point. In February 2026, NIST's National Cybersecurity Center of Excellence put out a concept paper specifically on software and AI agent identity and authorization. And the White House AI Action Plan from July 2025 named NIST across a significant share of its recommended policy actions.
Line those up and the trajectory is obvious: NIST is walking, deliberately, toward explicit agentic AI coverage. Banks building on the RMF now aren't guessing at a future standard; they're positioning themselves ahead of one.
The governance gap the RMF alone doesn't close — and what the CSA Agentic Profile adds
Here's the tension. NIST built the AI RMF for AI systems broadly. Agentic systems introduce risks the base framework never had in mind when it was written.
Three gaps stand out. Multi-step autonomous action with self-correction creates emergent behavior that the original framework doesn't explicitly account for. Tool-use and external API calls mean an agent's actions cross system boundaries, raising chain-of-authorization questions that a single-model framework was never built to answer. And delegation chains, where one agent instructs another, raise a genuinely hard question: when an orchestrating agent tells a sub-agent to act, who owns that outcome?
Deloitte's State of AI in the Enterprise 2026 found that only one in five companies has a mature model for governing autonomous AI agents, even as agentic AI adoption is set to climb sharply. That gap between adoption speed and governance maturity is the whole story in one data point.
The Cloud Security Alliance has proposed an answer: an Agentic Profile that extends RMF 1.0 function by function, adding concepts, categories, and subcategories for agent autonomy, runtime behavioral governance, and tool-use risk, without replacing anything underneath it. It's built to line up with CSA's own AI Controls Matrix, a 243-control framework across 18 domains published in July 2025, and with AAGATE, a reference architecture released in December 2025 that translates RMF principles into runtime governance for agentic systems specifically.
The practical takeaway: banks don't have to pick between NIST's framework and sector-specific controls. These pieces are meant to stack, one on top of the other. That stacked structure, GOVERN, MAP, MEASURE, MANAGE, plus the agentic extensions, is what the rest of this piece applies directly to banking.
GOVERN: building the institutional accountability structure before agents touch live transactions
GOVERN comes first for a reason. Without it, the other three functions have nowhere to live inside the organization.
For a bank deploying agentic AI, GOVERN means a few concrete things. AI risk needs to sit at the board level, treated with the same seriousness as any material operational risk rather than as an IT ticket. Roles need to be spelled out: who owns the agent's behavior day to day, who has authority to change what actions it's allowed to take, who reviews escalations when something goes sideways. And there needs to be an explicit, written policy on what an agent can do on its own versus what needs a human to sign off, including the dollar thresholds where that line sits.
Third-party vendor governance belongs here too. Since SR 26-2 puts model validation responsibility on the institution even for vendor-built systems, GOVERN is where that accountability gets written down and assigned to someone with a name.
The CSA Agentic Profile adds something specific worth flagging: delegation chain accountability as its own governance construct. Every agent in a multi-agent system needs a documented scope of authority and a named human who's accountable for it, at the individual level rather than the departmental one.
In practice, this often takes the shape of what you might call an AI agent charter: a document laying out permitted actions, prohibited actions, what triggers an escalation, and exactly who's responsible at each layer of the chain. Controls that are configurable and audit trails that are complete aren't features you bolt on after launch; they have to be built into the governance structure from day one, or they don't hold up under review.
Credit unions offer a useful, slightly uncomfortable snapshot here. Per Cornerstone Advisors' What's Going On in Banking 2026, 59% of credit unions have already deployed generative AI in some form. But fewer than 20% describe those deployments as enterprise-ready. That gap, adoption running well ahead of readiness, is almost certainly a GOVERN problem before it's anything else.
MAP: identifying where agentic risk concentrates in specific banking workflows
MAP's job is narrow on purpose: figure out the risk profile of this specific agent doing this specific task on these specific rails, rather than AI risk in the abstract.
A handful of banking workflows show where that risk concentrates hardest:
Payment and transfer execution, where an agent acts on a customer's instruction with no per-transaction human check, requires mapping out authorization boundaries, how fraud signals feed in, and velocity controls on repeat transactions. KYC and AML compliance work is another. One global bank reportedly deployed ten agent squads, each coordinating four to five AI agents, to move from periodic customer reviews to continuous, event-driven due diligence; MAP here has to account for how reliable each data source is, whether decisions can be explained, and whether that explanation would survive a regulator's questions at every handoff between agents.
Mortgage underwriting is its own case: agents pulling income, asset, and employment data autonomously from multiple sources, then simulating borrower risk, need MAP to trace exactly where each data point came from and flag any place regulatory cross-checks might be getting skipped. And collections is under real pressure right now; credit unions reported total delinquency of 95 basis points in Q3 2025, which is pushing institutions toward automated outreach faster than some may be ready for. MAP has to grapple with FDCPA limits on when and how an autonomous system can contact someone.
The 12 risks in NIST-AI-600-1 (prompt injection, hallucination, data poisoning, over-reliance, and others) give a starting checklist. CSA's extension adds tool-use risk mapping on top: every external API or data source an agent touches is a potential point of failure or attack, and it needs to be listed and rated, not assumed safe.
One more thing worth saying plainly: the map isn't a document you finish. Agent behavior drifts, new tools get added, use cases expand past what was originally scoped. MAP is something you keep doing, not something you filed once.
MEASURE: what meaningful monitoring looks like when an agent is executing transactions at speed
Here's the hard part. Agents make decisions faster than a person can review them, and the reasoning behind those decisions is often hard to see into. Sampling a handful of outputs after the fact, which is how a lot of traditional model monitoring works, just isn't enough anymore.
MEASURE for an agentic system in banking needs to include a few things at minimum. Audit logging needs to be immutable and near real-time, capturing not just the final transaction but the decision steps and tool calls that led there. There needs to be some way to detect behavioral drift, meaning the agent's pattern of actions shifting over time relative to what it's actually permitted to do. Any consequential financial decision needs an audit trail detailed enough to support an after-the-fact explanation that would hold up under regulatory examination. And for credit and collections specifically, bias and fairness need ongoing measurement against disparate impact thresholds, not a one-time check at launch.
Most institutions aren't there yet on infrastructure. Research across heavily regulated industries has found that most organizations lack dedicated AI incident reporting tools, even where AI risk has been formally identified. That's a striking number: it suggests most institutions have already identified the risks on paper but don't have the plumbing to catch those risks when they actually show up.
Governance architecture literature points to a fairly consistent technical baseline for doing this right: VPC isolation, immutable audit logs, controls over where data physically resides, role-based access, and policy enforcement sitting at the gateway layer where agent actions pass through. SOC 2 certification, an independent audit of security, availability, and confidentiality controls, gives a useful, verifiable floor to build that monitoring on.
MEASURE isn't the end of the chain, though. Without instrumentation, MANAGE turns reactive by default. You can't manage what you can't see.
MANAGE: operating agentic banking deployments with controls that hold under real conditions
MANAGE is where the loop closes: risk responses get put in place, tracked, and revised as conditions change. It's ongoing, day-in-day-out operational work rather than a milestone you hit and move past.
A few constructs matter most here. Human-in-the-loop thresholds set explicit triggers, transaction size, frequency, risk score, that route an agent's action to a person before it executes, not after. Kill switches and scope reduction give someone the ability to suspend or shrink what an agent's allowed to do instantly, without having to take down the systems around it. Exception escalation paths define exactly where an action goes when the agent hits the edge of its permission scope: who gets it, in what format, within what window of time. And incident response covers what happens when an agent does something it shouldn't have, including rollback capability, who gets notified, and whether the incident needs to be disclosed to a regulator.
CSA's Agentic Profile adds a runtime behavioral governance layer on top of all this: policy enforcement sitting at the gateway that can actually reject or modify an agent's action in the moment, against defined rules, rather than just logging what happened after the fact for someone to review later.
Where a bank deploys these controls matters too. Building agentic controls into a bank's existing infrastructure keeps the institutional governance structures already in place; standing up a brand-new, separate platform tends to mean rebuilding that governance from scratch, which is its own risk.
Jack Henry's expanded collaboration with Google Cloud, announced in June 2026 and serving roughly 7,400 community bank and credit union clients, is a decent real-world signal here: early adopters reportedly saw time savings of up to 70% on routine administrative tasks. What made that possible was a tightly defined operational scope established before anyone tried to scale it wider.
This is really where governance stops being a cost center and starts being a competitive edge. Banks that can show examiners a controlled, auditable agentic deployment tend to move faster, because they've earned confidence instead of triggering suspicion.
Putting the four functions together: what a phased agentic banking deployment looks like in practice
NIST built the RMF to be non-prescriptive on purpose; it doesn't tell you the order to do things in. Still, a phased rollout maps onto the four functions pretty naturally, and it's worth walking through what that looks like.
Phase one is GOVERN, done before an agent ever touches live data: write the agent charter, assign accountability, sort out third-party validation responsibilities so there's no ambiguity about who owns what once something's live.
Phase two is MAP applied to a single use case, chosen deliberately small and bounded, something like a payment status inquiry or a first pass at collections outreach, where the blast radius of a mistake is limited and the workflow is well understood.
From there, MEASURE gets built around that one use case, with real audit logging and drift detection running before volume scales up, and MANAGE gets tested with actual human-in-the-loop thresholds and a kill switch that's been proven to work, not just documented. Only once that first narrow use case runs clean does it make sense to widen the scope. That's the pattern that separates the 9 institutions actually running agents in production from the much longer list still stuck in pilot purgatory: a governance sequence built before the agent went live, rather than bolted on after something went wrong.


