Rail Governance

Board-Level AI Governance Structures for Transaction-Executing Systems

Banks are rushing to deploy AI agents while governance structures lag dangerously behind.

Senior Writer · · 12 min read
Cover illustration for “Board-Level AI Governance Structures for Transaction-Executing Systems”
AI Governance Frameworks · September 15, 2026 · 12 min read · 2,760 words

Agentic AI doesn't recommend or flag something for a human to decide. It reads context, plans a sequence of steps, and executes transactions on its own, no continuous prompting required. That shift, from advising to acting, is what's forcing a new tier of board-level governance inside banks and credit unions, one built around execution authority rather than model accuracy alone.

A model that scores a loan application needs oversight of its outputs. Was the score fair? Was it accurate? An agent that starts a wire transfer needs something different: oversight of authority. Who said this agent could move money? Under what conditions? What stops it from acting outside those bounds? Existing model risk frameworks, SR 11-07 and the NIST AI RMF among them, were built for a different generation of AI systems. They weren't built for chains where one model's output feeds straight into another model's action, closing a financial loop with no human anywhere in the execution path.

Most banks still bolt governance on after the model goes live, checking outputs once they're already in production instead of setting boundaries before the system runs. That sequence worked fine when a model just scored something and a person decided what to do with the score. It stops working once an agent operates across systems, channels, and other agents at scale, because by the time anyone reviews the output, the transaction has already happened. Cisco's AI Readiness Index for 2025 puts a number on the gap: only 31% of organizations say they're fully equipped to control and secure agentic AI systems, even as 83% plan to deploy them anyway. Boards are being asked to sign off on a category of risk most institutions admit they can't fully contain yet.

The regulatory environment that is pushing this to the boardroom now

Diagram: The Governance-Deployment Gap in Agentic AI. Visualizes: Visualize the stark gap between two statistics from Cisco's AI Readiness Index 2025: 83% of organizations plan to deploy agentic AI systems, yet only 31% say they are fully equipped…

Treasury didn't wait around. On February 19, 2026, the department released the Financial Services AI Risk Management Framework, the first federal, sector-specific AI risk resource built for the financial industry. More than 100 institutions had input into it. The framework translates the NIST AI RMF into 230 control objectives spread across four adoption stages: Initial, Minimal, Evolving, and Embedded.

It's voluntary. But voluntary doesn't mean optional once an examiner walks in the door asking what "reasonable" looks like, because this is now the document they'll point to.

State legislatures haven't slowed down either. The NCSL's AI Legislation Database logged 1,208 AI bills introduced across all 50 states in 2025, with 145 enacted. Colorado's AI Accountability Act, taking effect January 1, 2027, is the most far-reaching state law on the books right now, focused on transparency and requiring that consumers be told when AI is involved in matters affecting them.

Overseas, BaFin's December 2025 guidance treats AI as an operational risk topic, full stop. It's ICT risk, governed under DORA. The OCC's Fall 2025 Semiannual Risk Perspective flagged AI risk explicitly too, flagging it alongside other operational and technology risk categories.

None of the older statutes take a pause for AI, either. ECOA, FCRA, CFPA, GLBA, BSA: they all apply whether the actor making the decision is a person or an agent. Existing data-protection requirements around access controls and monitoring don't disappear because an AI agent is the one touching customer data.

EY's May 2025 Responsible AI Pulse Survey found 52% of banks naming governance as their top challenge in adopting AI, yet only 44% said they were actually directing investment toward technology that supports ethical AI use. That's a real gap between what banks say worries them and what they're building to fix it. EY's October 2025 follow-up survey found 98% of respondents had already taken financial losses tied to AI-related risks. Boards can't file this under "emerging" anymore. It's already cost money, a fact worth sitting with.

How board-level AI governance has actually been restructuring inside financial institutions

Something concrete is happening at the board level, not just talk. A BPI/BITS survey of large, complex financial institutions (GSIBs, super-regional banks, trust and asset servicing firms, insurers) found that 64% have already formed, or are actively weighing, a dedicated board technology committee. Of those, 94% set it up as a standing committee of the board itself, not a subcommittee tucked under audit or risk. That's a structural change, not a paperwork exercise.

Why form the committee in the first place? Focused oversight of technology strategy and major investments led the reasons, cited by 81% of respondents. Support for digital transformation and emerging tech, including AI, came in at 56%. Cybersecurity and operational resilience oversight followed at 38%. Per BPI, AI and emerging tech were among the key topics boards focused on in 2025, including GenAI and agentic AI controls.

Creatio's Global AI and No-Code Adoption Survey backs this up from another angle: 80% of financial services leaders say AI agents are already a board-level topic, or will be one in 2026. But forming a committee and knowing what that committee should actually govern are two different problems, and most institutions are further along on the first than the second. EY's October 2025 survey found institutions still sorting out which executives own AI governance and which risk indicators even apply.

McKinsey's 2026 AI Trust Maturity Survey puts a number on that lag: only about a third of organizations report mature governance across strategy, oversight, and agentic AI specifically. Financial services leads most other sectors on adoption speed. Governance maturity hasn't caught up, and the gap between the two is where the exposure sits.

What's being asked of boards has changed in kind, not just in volume. Understanding what the technology does isn't enough anymore. Directors are expected to push management directly on risk appetite, on where execution authority should stop, and on whether the money spent on AI is actually producing value that matches the risk it brings in.

What the three-lines-of-defense model covers and where it stops when AI executes transactions

The three-lines model still does real work here. Business units own the risk they create. Risk and compliance teams oversee the controls. Internal audit checks independently that the first two are actually doing their jobs. That structure holds for AI the same way it holds for credit underwriting or fraud review.

For analytical AI, the standard model risk management toolkit still fits: model inventories, validation before deployment, bias and fairness testing, ongoing monitoring for drift. None of that needs reinventing.

But watch what happens once AI stops analyzing and starts acting. Traditional MRM assumes a human sits somewhere in the loop before an action completes. An agent that executes a transaction removes that assumption entirely, and the framework was never built to notice the assumption is gone. Existing model risk guidance wasn't written with autonomous agents in mind and offers no clear framework for systems that plan a sequence and carry it out without stopping for review.

In a multi-agent setup, the risk compounds. What matters is whether one model fails on its own. It's what happens when several agents act in sequence and their combined behavior produces something none of them would have produced alone. A governance failure at step two of a five-step chain might not surface until the transaction has already closed and the money has already moved.

So where's the actual line? Integration connects systems and moves data around. Orchestration coordinates several capabilities into a business process. Execution is the layer where an agent takes a committed action on a real account: real dollars, real consequence. Governance built only for the upstream pipeline, the integration and orchestration layers, misses the layer that actually matters. Most boards, asked directly whether their current structure assigns clear accountability for that execution layer, would have to admit it doesn't.

The four governance components that a transaction-executing AI system specifically requires

Diagram: Four Required Governance Components for Transaction-Executing AI. Visualizes: Show the four non-optional governance components that a transaction-executing AI agent specifically requires, in the order they must operate: (1) Execution…

Four pieces here, and none of them optional once an agent is touching real transactions.

An execution authority registry. Every agent in production needs a documented, board-approved scope: which transaction types it can start, up to what dollar value, under what conditions, and what triggers an automatic halt or escalation. This isn't the same document as a model inventory. A model inventory describes what a model does. This registry describes what it's allowed to do, a different question entirely, and the one that matters most to a regulator.

Real-time permission enforcement at the point of execution. Authority has to be checked the moment an action happens, not attested to in a policy binder somewhere. Backbase's Sentinel layer, part of its AI-native Banking OS launched in April 2026, is one illustration of what this looks like structurally: it checks an actor's permissions against bank policy before the action runs, and logs it. Whatever platform a board is looking at, that same real-time check needs to exist somewhere in the stack.

An audit trail that reconstructs the decision, not just the outcome. A transaction log showing what happened isn't enough for an examiner. The record needs to show what data the agent pulled, which model version produced the output, what policy governed the step, and who, or what, authorized the specific action taken.

Escalation and halt architecture with thresholds the board itself sets. Configurable rules define when an agent must stop and wait for a human, what categories of action it's never allowed to take on its own, and who has the authority to override or shut the whole thing down. Those thresholds are a governance decision, not something quietly handed off to the engineering team.

EY frames this as a "watchtower" model: clear oversight lines at board and executive level, spanning the full AI lifecycle, with technology, risk, legal, and other functions all in the room. For systems executing transactions, the watchtower's job is real-time visibility into what's happening at the execution layer right now, not a retrospective report six weeks later. Gartner's research on this is blunt: organizations that deploy AI governance platforms are 3.4 times more likely to hit high effectiveness in AI governance. Governance designed into the system enforces itself. Governance added on top after launch generally doesn't, and that difference is the whole ballgame.

How the governance structure connects to the board's accountability for specific transaction risk classes

Payments and transfers sit at the sharpest edge of this. A wire, once sent, is irreversible. The governance question asks whether the underlying model is accurate on average. It's whether this specific execution, at this specific moment, fell within the authority the agent had actually been given.

Fraud detection carries its own version of the same tension. An AI system flagging and freezing an account in real time is making a consequential call either way: fraud that slips through costs money, a legitimate transaction wrongly blocked costs trust. Wolters Kluwer's Q1 2026 Banking Compliance AI Trend Report found explainability and transparency cited as the most acute regulatory concern by financial institutions, at 28.4%, with bias and discrimination close behind.

Credit decisioning isn't a new fair-lending concern; regulators have watched this territory for decades. But agentic systems that approve or deny credit on their own shrink the window a human reviewer used to have to catch an edge case before it became a decision on the record.

AML and KYC workflows sit squarely in BSA and USA PATRIOT Act territory. An agent running transaction monitoring or identity verification has to produce evidence of the same quality a human reviewer would produce, because the legal standard doesn't relax just because software did the work.

Customer-facing agents handling account changes, payment scheduling, balance transfers, need to operate under the same boundaries a human customer service rep would. The FTC Safeguards Rule's access controls apply whether the hands touching customer financial data are human or not.

There's a systemic layer above all of this that's easy to miss when thinking institution by institution. There is a systemic concern that runs above the individual institution level: when many agents across many institutions respond to the same signal at once, their combined behavior could produce outcomes well beyond what any single institution intended. Board governance has to account for what one agent does, and what happens when many agents, at many institutions, react to the same signal at once.

What credit unions need to build differently given leaner governance resources

Credit unions carry the same compliance load as the largest banks. NCUA readiness, BSA, ECOA, data protection: none of that scales down just because the institution is smaller.

What does scale down is the governance apparatus available to meet those demands. Boards at community institutions typically don't have a dedicated technology committee, and they may not have a single director with deep AI or cybersecurity background. The BPI/BITS survey cited earlier drew its 64% figure from GSIBs, super-regional banks, trust and asset servicing firms, and insurers, so that number simply doesn't describe the credit union sector. Copying the big-bank committee model at a fraction of the staff isn't a plan, it's a way to fall behind slower.

In practice, credit unions need the same four governance components (execution authority documentation, real-time enforcement, complete audit trails, escalation thresholds) with a fraction of the dedicated staff and a much smaller internal audit function to run them. A few moves fit that reality better than trying to shrink a big-bank structure down to size:

  • Fold the execution authority registry and model inventory into one governed document, owned by a named officer, reviewed by the board on a set schedule, instead of maintaining separate systems nobody has time to reconcile.
  • Insist that any agentic platform brought in comes with native auditability and configurable controls built in. A platform that makes the credit union build its own governance layer on top isn't a fit, no matter how good the underlying model is.
  • Use Treasury's Financial Services AI Risk Management Framework, released February 19, 2026, as the self-assessment tool of choice. Working directly from the document, mapped to 230 control objectives across four maturity stages, the same one examiners will reach for, is the shortest path to being ready when they do.

The right tooling can make a meaningful difference at community-institution scale. Conversation-level guardrails, built into the live execution path rather than reviewed after the fact, brought the institution closer to NCUA compliance readiness, at an institution navigating significant internal uncertainty about how to proceed with generative AI.

Some credit unions have pursued agentic platforms to automate routing and recommendations across a wide range of case types. The governance lesson there is subtle but important: the platform's own structure (centralized data, clearly defined case types, routing logic that leaves a trail) does governance work the credit union would otherwise have had to build from scratch.

Either way, the board's job doesn't change. Set the execution authority thresholds, and demand proof the platform actually enforces them. The technical build can come from a vendor. The decision about what an agent is allowed to do cannot.

How to evaluate whether an agentic AI platform is compatible with board-level governance requirements

One test cuts through most of the vendor noise: does the platform check execution authority at the moment an agent acts, or does it just log what already happened and call that governance? A system that records the action after the fact is a fundamentally different thing from one that checks permission before the action runs. Boards should require the second kind and treat the first as a gap, not a feature, no matter how the sales deck frames it.

Backbase's AI-native Banking OS, launched in April 2026, is one example of a platform built around that distinction from the start. It's structured in layers, including an Intelligence layer that reads signals, a Semantic Layer, with a component called Nexus, and Sentinel, the authority layer that checks every actor's permissions against bank policy before it executes anything, logging the action as it happens. The platform serves more than 120 financial institutions, including Navy Federal Credit Union, TD Bank, and KeyBank. Backbase also folded in banking-tuned conversational agents through an earlier acquisition, building that capability directly into the same governed structure instead of bolting it on separately.

Walking through an example like this is meant to make the underlying test concrete, not to crown one vendor over the field. Ask any platform under consideration to show, specifically, where in its architecture an agent's permission gets checked before it acts, and what gets written down when it does. If a vendor can't answer that with a specific mechanism, the answer probably doesn't exist yet, and the board is the last line of defense standing between that gap and a transaction that shouldn't have happened.

Sources

  1. Moving from pilot to production | Wolters Kluwer
  2. Cyber Board Governance: The Role of Board Technology Committees for Financial Services Companies - Bank Policy Institute
  3. Five hallmarks of effective AI strategies in banking
  4. AI Governance Is Now a Board Responsibility

More in AI Governance Frameworks