How Bank Examiners Test AI Transaction Controls in Safety and Soundness Reviews
Regulators expect banks to document AI controls now, even without final rules.

No AI examiner rulebook exists yet, and that's exactly the problem. Bank examiners test AI transaction controls by pointing existing safety-and-soundness, model-risk, and compliance frameworks at whatever new system a bank has bolted onto its payment rails. Reuters reporting citing three people with direct knowledge says the OCC and Federal Reserve have made AI a standing item on every periodic exam now, and the OCC has already folded AI findings into enforcement actions in 17 matters since fiscal year 2020. Those are real consequences already on the books, and the OCC said as much in its Fall 2023 Risk Perspective: new technology doesn't retire old safety and soundness standards. So the question sitting in front of every bank leader isn't whether the rules are settled. What matters is whether you can describe your own controls to an examiner right now, in the room, without hedging.
How the regulatory framework shifted in April 2026, and what that leaves unresolved
SR 26-2 from the Federal Reserve and OCC Bulletin 2026-13 took effect April 17, 2026, replacing SR 11-7, the model-risk guidance that had shaped bank governance since 2011. Fifteen years is a long run for any regulatory document. The replacement was overdue by most measures, and most of the industry saw it coming.
Here's the catch, though: the new guidance carves generative and agentic AI out of scope entirely. Regulators decided those systems move too fast and behave too unpredictably to lock into a fixed rulebook right now. An RFI on model risk management is supposedly coming, one that will deal with generative and agentic AI head-on, but it hasn't been issued yet. That leaves a real gap between what's written down and what's actually running inside banks today.
Does that gap mean banks get a pass? Supervisors aren't treating it that way. Institutions are still expected to apply model-risk principles to any AI system that matters, rule or no rule. The bar examiners hold banks to looks almost identical to what SR 11-7 always demanded: documented development and use, independent challenge, clear governance ownership, and an audit trail an examiner can actually sit down and inspect.
Credit unions face a parallel version of this problem, and arguably a worse one. Supervisory reviews have found that NCUA lacks both comprehensive model-risk guidance and the authority to directly examine third-party AI providers. Credit unions are getting measured against a standard that hasn't been formally written down, while their exposure to third-party AI vendors keeps growing underneath them. The absence of a rulebook doesn't shrink examiner scrutiny. It just hands examiners wide discretion to reach for whichever existing framework fits the risk sitting in front of them.
The specific control dimensions examiners probe first
Four lines of inquiry show up again and again in supervisory reporting. Know them cold before an examiner walks in, because "we're working on it" doesn't hold up as an answer.
Governance and ownership come first. Can the institution name, specifically, who is accountable for each AI system in use? Is there board-level visibility into how these systems actually work, not just a slide deck handed down from IT? The OCC has said plainly that generative AI falls under safety-and-soundness expectations, so board-level oversight isn't optional there.
Second: technical limits on model behavior. What actually stops the AI from doing something it shouldn't, particularly around transaction execution? Examiners want these guardrails written down and demonstrable, not described in a hallway conversation on exam day.
Third, human review structure. Where, exactly, does a person enter the loop on an AI-driven transaction? Examiners look for documented escalation paths and specific dollar or risk thresholds. A vague assurance that "humans are involved somewhere" doesn't survive five minutes of questioning.
Fourth: kill-switch and emergency shutdown capability. Can the institution actually turn an AI system off, fast, if something goes sideways? A Wolters Kluwer survey of banking professionals found that 72% flagged kill-switch protocols as a gap in their own governance. That's one of the most common weak spots examiners are finding in the field right now, and it's arguably the easiest one to fix, which makes it the hardest to excuse.
A fifth thread runs underneath all of this: data boundary enforcement. Examiners ask, point blank, whether AI tools are reaching into data they were never authorized to touch. That question gets sharper the more AI models get built to pull from multiple data sources at once, which raises privacy and compliance issues that didn't exist in the same form five years ago. Exam rooms are asking these questions right now, this cycle.
How transaction testing works when AI is executing the transactions
Transaction testing has always been a core examination technique, the method examiners use to check that a bank actually follows its own stated policies. AI complicates that step considerably.
Here's the meaningful shift: examiners calibrate how deeply they probe based on the quality of an institution's own monitoring and the risk profile of its AI systems. Run a rigorous, well-documented independent testing program, and face less direct examiner testing as a result. Run a weak or undocumented one, and invite a much closer look. The quality of a bank's internal testing shapes how hard examiners dig on their own, and that's the trade sitting on the table.
What does "well-documented" actually mean once AI is in the loop? Sound practice requires that any new model version be tested against prior outcomes before going live, with discrepancies documented and resolved. That keeps orchestration decisions replicable, something an examiner can walk back through step by step even after several model updates have come and gone.
Fraud detection adds another layer. AI-driven systems were reportedly catching 92% of fraudulent activity before approval as of late 2025, according to ABA Banking Journal, a strong number on its face. Examiners aren't going to take that figure at face value, though. They'll want to see how the institution validates and audits that detection performance over time, not just the headline stat. An institution's own testing program is the primary artifact examiners use to decide how hard they need to push.
What an audit trail must show when AI initiates or executes a payment
Every AI-driven transaction needs to trace back to a few basic facts: who authorized it, what parameters governed the decision, what the system actually decided and why, and whether a human reviewed it, or at least could have, before it went out the door.
Agentic AI changes this in a specific way, and it's worth sitting with. A traditional model produces a recommendation, and a person acts on it. The decision and the action are two separate events, split apart by a human standing in between. An agentic system that executes a payment collapses that separation entirely: the decision and the action happen at the same moment, by the same system. So the audit trail has to capture the whole chain, start to finish, not just the final output.
Regulators and standards bodies have been fairly direct about what they expect here. Regulators expect full logging as a baseline for AI running in regulated environments. Sound practice calls for secure development lifecycles alongside full logging and ongoing adversarial testing, the kind that actively tries to break the system rather than just checking it works under normal conditions. OCC Bulletin 2026-13 lists outcomes analysis and clear governance, with defined policies and roles, as features of sound model use. Translate that into an agentic context, and each payment execution becomes its own outcome, one that needs to be logged against whatever policy was supposed to govern it.
Third-party AI adds a layer a lot of institutions haven't fully reckoned with yet. Industry surveys have found that a significant share of firms report only partial understanding of the AI technologies they are actually using. Examiners expect the bank to own the audit trail regardless of who built the underlying model. That means institutions using vendor-delivered AI for payments need real contractual and technical access to the logs an examiner would need to see. "The vendor handles it" doesn't hold up as an answer in an exam room.
The governance gap that makes most institutions exam-vulnerable right now
Look at the size of the mismatch. Seventy percent of banking firms are already using agentic AI in some capacity, but very few describe their governance strategy as well-defined and properly resourced, according to a 2026 industry report. Most institutions, in other words, are already being examined on systems they haven't finished governing.
Credit unions sit earlier on the curve, with 17% having deployed agentic AI according to Cornerstone's 2026 report, a smaller share. But Supervisory expectations around AI and third-party risk are already shaping credit union examinations, so even early adopters aren't getting a grace period.
The AI inventory problem sits underneath most of this. Examiners expect a documented list of every AI-enabled system in use: its purpose, the data it touches, who owns oversight of it. Plenty of institutions can't produce that list, especially for vendor tools that got adopted informally, one department at a time, without ever going through a formal review. PaymanAI, for instance, builds that inventory assumption in from the start, running AI agents for banks on existing rails with configurable controls and a full audit trail behind every transaction. According to a 2026 report from the Cambridge Centre for Alternative Finance, between 70% and 74% of firms integrating AI cite data privacy exposure, hallucination, and lack of explainability as top concerns. An institution that can't explain its own AI systems to itself has no shot at explaining them to an examiner.
BSA/AML monitoring carries its own specific exposure. AI used in anti-money-laundering work has to be auditable, free of bias, and defensible on its face, and regulators have made clear that opacity isn't acceptable here even when the underlying detection performs well. A model that catches suspicious activity but can't explain how it got there is still a problem, whatever its detection rate says.
This governance gap is operational, and it loops back into the testing discretion point from earlier. Institutions without clear, documented governance can't take advantage of examiners' willingness to lean on internal testing, because that discretion only kicks in when the testing itself is qualified, independent, and matched to actual risk. No governance means there's nothing for that discretion to attach to.
What exam-ready AI governance looks like in practice for banks and credit unions deploying agentic systems
Strip all of this down, and a documented governance framework needs a few concrete pieces in place before an examiner ever asks for them. Skip any one of these, and the rest of the framework is decoration.
Named accountability at the board level, not something quietly handed off to technology staff. A current inventory of every AI system in use: its function, what data feeds it, who owns it. Guardrails documented and enforced technically, not just written into a policy binder somewhere. A human-review structure with defined escalation thresholds, particularly for transaction execution above set risk levels. A kill-switch that's actually been tested, not just theoretically available if someone remembers where it is. Audit logs an examiner can get into for every AI-initiated transaction. An independent testing program that triggers regression testing automatically whenever a model or policy changes.
For agentic systems specifically, governance has to cover what the AI executes, not just what it recommends. The decision and the action are the same event, and both need to be auditable together, not separately.
There's a practical edge worth naming here too. Deploying AI on existing banking rails, rather than standing up entirely new infrastructure, shrinks the scope of what needs governing and keeps audit trails in one place instead of scattered across systems. That matters when an examiner asks how the AI actually interfaces with core banking systems, because a clean answer beats a complicated one every time. SOC 2 certification and general compliance-readiness aren't optional credentials anymore either. They're the baseline that lets an institution credibly tell an examiner its vendor relationships have been checked against a recognized standard.
Well-documented transaction testing matters even more once AI agents are the ones executing the transactions, since institutions have to show consistent, auditable results across repeated runs, not a good outcome once and a hope that it holds. Platforms designed to execute real banking transactions through AI agents should build in full audit trails and configurable controls specifically so they hold up under that kind of scrutiny, giving examiners visibility into how each transaction was authorized and whether outcomes stay consistent across similar scenarios.
The institutions in the best shape when an examiner shows up are the ones that treated governance as a design constraint from day one, built before the system went live and started running transactions, not bolted on afterward because an exam forced the issue. Configurable controls, complete audit trails, visible human oversight: these work best built into the architecture from the start, before an exam is even on the calendar. The exam is the accountability mechanism that makes adoption stick. Institutions that build to that standard now are the ones that get to keep expanding what AI does for them, without a regulator stepping in to slow it down.


