Rail Governance

Cybersecurity Examination Expectations for AI Systems on Banking Rails

Regulators now demand proof of AI governance, not just compliance checkboxes.

Staff Writer · · 11 min read
Cover illustration for “Cybersecurity Examination Expectations for AI Systems on Banking Rails”
Examiner Expectations · September 11, 2026 · 11 min read · 2,457 words

Cybersecurity exams for AI on banking rails have stopped being a checklist exercise. Examiners now spend their time asking who owns a given AI system, how decisions get documented, and whether the bank can defend its governance in the room. That shift changes what "passing" actually means.

The regulatory framework that governs AI on banking rails right now

On April 17, 2026, the OCC, the Federal Reserve, and the FDIC issued revised model risk management guidance, known as SR 26-2, replacing the 2011 framework that had governed model risk for over a decade. The new guidance is principles-based and gives institutions more room to shape their own programs.

Here's the catch. SR 26-2 explicitly leaves generative and agentic AI outside its formal scope. The agencies call these systems novel and fast-moving, and they've said a request for information on how banks use AI, including agentic systems, is coming. That guidance doesn't exist yet.

Don't read the exclusion as relief. It moves the burden of figuring out governance onto the bank, not off of it. Supervisors still expect model-risk thinking applied wherever AI touches something that matters: an agent that influences a credit decision, flags fraud, or feeds financial reporting needs documented governance, an inventory entry, monitoring, and evidence a bank can hand to an examiner on request. Nobody's letting you skip validation just because the word "agentic" doesn't appear in the rule.

Other pieces of the framework are moving too. NIST put out a preliminary draft of its Cybersecurity Framework Profile for Artificial Intelligence in December 2025, organized around three areas: locking down AI system components, using AI for cyber defense, and stopping AI-enabled attacks. It's meant to sit alongside NIST's existing AI Risk Management Framework, not replace it, and the Cybersecurity Coalition has pushed NIST to treat AI as software rather than invent an entirely separate compliance universe. The NCUA has pointed to the NIST AI Risk Management Framework as a relevant resource for credit unions, meaning credit unions face broadly similar governance expectations around AI risk management.

Look overseas and the picture sharpens. DORA, in force across the EU since January 17, 2025, requires continuous monitoring of ICT systems with named management accountability, plus logging and classifying incidents, including ones tied to AI. The European Central Bank said in February 2026 that banks need clear accountability for AI-driven decisions and real challenge mechanisms involving risk, compliance, and internal audit, warning that AI projects without a strategic anchor produce fragmented governance and risks nobody sees coming. US expectations aren't there yet, but DORA reads like a preview.

Meanwhile the SEC's FY2026 exam priorities put cybersecurity and AI ahead of cryptocurrency as the top concern, with examiners expected to dig into AI in client-facing tools and demand documentation of testing and ongoing monitoring. The direction across every one of these bodies points the same way: less prescription, more proof.

What agentic AI actually does on banking rails and why it creates a distinct risk profile

An agentic system doesn't just suggest something and wait. It acts. On banking rails, that means an agent can initiate a payment, approve an exception, reroute a transaction, trigger a compliance check, manage liquidity inside real-time gross settlement systems, or run a cross-border payment chain start to finish.

None of that would be technically possible at this scale without the ISO 20022 migration, which gave payment messages richer, structured, machine-readable data across the entire lifecycle. That data layer is what makes an agent's autonomy usable rather than theoretical.

So what happens when the agent gets it wrong? A single bad answer from a chatbot is a bad answer. A single bad decision from an agent with execution rights is a transaction, and at scale it's a lot of transactions before anyone notices the pattern. That's the real distinction: the risk isn't one wrong answer, it's a cascade of wrong actions moving faster than a human review cycle.

The deployment numbers back up how fast this is moving. McKinsey counted more than 160 AI use cases announced across 50 of the world's largest banks in 2025 alone. Yet actual production deployment lags way behind the announcements. Research from Ness Digital Engineering found nearly all companies plan to put agents into production, but only 9 to 14% actually have. Separate research put the share actively running agents in production at 11%, and governance and security concerns, not technical capability, are widely cited as the reason for the gap.

Why the gap? Ask the banks themselves. The same research found 48% of organizations point to governance concerns as their most pressing challenge, and 20% say their own data isn't ready to support it. Across industry research, banking executives consistently point to governance, risk, and compliance as the primary barriers between them and realizing value from agentic AI.

Here's the mental model that actually fits: an agent that executes a transaction isn't like a model that scores a risk. It's closer to an employee who has system access. Employees get permission structures, activity logs, and a manager who signs off on what they're allowed to touch. Agents need the same thing, and examiners are starting to ask for it in those terms.

The specific threat vectors that examiners will probe for agentic systems

Banking ML systems face a range of well-documented attack types, including data poisoning, model extraction, and evasion attacks. These attack types are well documented in AI security research, so this isn't hypothetical.

Prompt injection sits at the top of OWASP's list of LLM risks, and in an agentic banking setting the stakes change completely. A misleading chatbot answer is annoying. A manipulated agent that moves funds, approves an exception it shouldn't, or exposes account data because of an injected instruction is a breach.

Integration middleware deserves particular attention here. A layer sitting between the AI agent and core banking systems can act as a concentration point, meaning one vulnerable component can potentially reach several core systems at once instead of just one. Concentration risk, in other words, dressed up as convenience.

The attacker's toolkit is getting sharper too. Readily available tools lower the bar for sophisticated social engineering, and deepfake voice cloning is already being used to trick bank employees into authorizing wire transfers that never should have gone out. Layer that onto faster payment rails and the math gets worse: the faster money moves, the less time anyone has to check whether it's going to the right place. Duplicate payments slip through. Scams complete before the warning signs even surface.

None of this is abstract to the sector holding the money. Kings Research put the BFSI industry's share of the AI infrastructure security market at 42.03% in 2025, which tells you where the perceived exposure is concentrated. And KPMG found that 74% of financial services firms say cybersecurity gets involved from the earliest planning stages of a new technology investment. Good sign, except involvement at the planning table doesn't equal a documented control at execution. Planning meetings don't stop a bad model update from shipping.

Third-party AI models raise their own version of the same problem: NIST's draft flags the need to verify integrity of vendor models to catch backdoors and supply chain compromises before they're baked into production. Put simply, examiners aren't asking whether the bank ran a penetration test. They're asking whether the threat model got written down, who owns it, and whether it maps back to the bank's AI inventory.

What documented controls and clear ownership look like in practice

Accountability that lives in someone's head doesn't count for much in an exam room. It has to show up in committee minutes and board packages, the actual paper trail examiners read, and DefenseStorm's observation is straightforward: that record needs quality data behind it, data good enough to support real decisions and real challenge, not a rubber stamp.

Getting AI-related decisions written into board minutes, with a name attached, before any agent goes live touching a customer or a member, might be the single highest-leverage governance step available to a bank right now.

Cornerstone Advisors found agentic AI is already a board or executive-level topic at more than half of financial institutions surveyed. Progress, sure, but there's still a wide gap between "we discussed it" and "we have a signed policy." Supervisory expectations for credit unions broadly mirror what examiners look for elsewhere: a board-approved AI policy, a named owner accountable for it, and documented incident response procedures. Three things, none of them exotic, all of them frequently missing.

Every AI system that matters, including any agent that can execute a transaction, needs to sit in an inventory with an owner attached and its governance written down. That holds even though SR 26-2 formally leaves agentic systems out of its scope. Supervisors expect testing records, monitoring evidence, and a clear answer to "what does this system do and who's allowed to change it," validation requirement or not.

Permission boundaries are where a lot of this becomes concrete rather than aspirational. Can the bank set limits on what an agent is allowed to do, to whom, and up to what dollar threshold? Can it log every single action against those limits, not just the ones that trip an alert? That configurability is the difference between a system that's merely running and one that's actually governed.

Human oversight still has a seat at the table too. Consumer protection and fair banking principles call for clear escalation paths when an AI-related complaint comes in, and human review for decisions that affect a customer's finances in a meaningful way. Autonomy doesn't get to replace accountability, full stop.

And third-party AI risk isn't a diligence checkbox, it's a governance question in its own right. Examiners want to know whether those third-party and fourth-party risks are actually managed, not just whether a vendor questionnaire got filed. A SOC 2 report is a real signal of operational discipline from a vendor touching banking rails, worth requiring, and worth actually reading rather than filing away unread.

How vendors currently deploying agentic AI in banking have built governance into the architecture

Backbase launched its AI-native Banking OS in April 2026, an operating layer that sits above the core banking system, payments rails, and CRM. It's built in three layers. An intelligence layer handles signals. A Nexus layer holds a shared semantic record that every agent and every employee reads from, so nobody's working off a stale or conflicting version of customer data. And a Sentinel layer checks every actor's permissions against bank policy before anything executes, then logs the action regardless of outcome. Backbase counts Navy Federal Credit Union, TD Bank, and KeyBank among its named clients.

Eltropy, which serves more than 750 community financial institutions, launched a governed agentic AI platform for credit unions in March 2026, built specifically for the governance realities smaller institutions face. Informed.IQ, meanwhile, has built what's described as the most advanced agentic implementation currently running in lending verification: vertical AI agents trained on a proprietary dataset covering more than 2 billion data points drawn from over 100 million loan documents. The company raised $63 million in December 2025, led by Invictus Growth Partners.

Different products, same architectural instinct. In each case, a permission and logging layer sits between the agent and the core banking system. Nothing executes until a policy check clears it, and every action gets logged whether it succeeds, fails, or gets kicked back for review.

That pattern maps almost directly onto what examiners are asking for: a named permission boundary, a log that functions as an audit trail, and a policy document the log can actually be checked against. Building the governance layer on top of existing rails, instead of replacing the core system outright, means the bank's existing governance structures stay intact while the automation runs through them. The governance travels with the agent. It doesn't live somewhere separate from it.

None of this closes the gap on its own, though. Cornerstone's Data EQ study, which surveyed 124 banks and credit unions in 2025, found an average data readiness score well below the midpoint among community financial institutions. A governed architecture only works as well as the data feeding it, and a lot of institutions aren't there yet.

The audit trail as the exam artifact that ties controls to accountability

A log file is not an audit trail. An audit trail connects a specific action an agent took to the specific policy that allowed it, the specific permission that governed it, the specific person accountable for it, and the exact moment it happened. Miss any one of those links and what you're holding is just data, not evidence.

Examiners reading board minutes want proof that AI-driven decisions got surfaced to leadership with data good enough to support a real challenge, not a summary nobody argued with. A log nobody reviewed isn't an audit trail. It's a liability sitting in a folder waiting to be found.

A defensible audit trail for an agentic system needs to answer a specific set of questions: what was the agent authorized to do, and who authorized it? What did it actually do, with timestamps attached? What controls got checked before execution happened? What exceptions came up, and how were they escalated? And who reviewed the log, and when?

The SEC's FY2026 exam priorities spell this out directly, expecting firms to keep detailed documentation covering both AI testing protocols and the results of ongoing monitoring. Pre-deployment and post-deployment both need a paper trail, not just one or the other. DORA's incident classification requirements for ICT-linked events offer a working model of what that documentation looks like in practice, and it's a reasonable bet that US expectations drift in that direction over time.

The OCC's Matters Requiring Attention process is the mechanism that turns weak AI governance into a formal supervisory finding. An audit trail the bank can't produce, or can't explain when asked, is about the fastest route to an MRA that exists.

So what actually separates the banks that walk out of an exam meeting early from the ones that don't? It's the ones that can walk an examiner through one specific agent action, start to finish. The policy that authorized it. The permission check it passed through. The log entry it left behind. The board report that referenced it later. That's not a compliance exercise, it's a demonstration that the system is owned, understood, and explainable end to end.

The real question examiners are asking in 2026 was never whether a bank should deploy agentic AI on its rails. It's whether that deployment can be explained, owned, and defended when someone from the OCC sits down across the table and asks. The audit trail is the answer to that question, in exam form.

Sources

  1. Cybersecurity Considerations 2025: financial services
  2. What Bank Examiners Are Prioritizing in 2026: Cyber Edition
  3. ECB AI Cybersecurity Letter to Banks: Quantum Is Next
  4. occ.gov
  5. sullcrom.com
  6. U.S. bank regulators scrutinize AI use in routine bank exams
  7. neurons-lab.com
  8. backbase.com

More in Examiner Expectations