Rail Governance
FeaturesLong read

Examiner Findings in Published AI-Related Enforcement Actions at Banks

Regulators are holding banks accountable for AI systems using decades-old compliance standards.

Senior Writer · · 11 min read
Cover illustration for “Examiner Findings in Published AI-Related Enforcement Actions at Banks”
Features · August 28, 2026 · 11 min read · 2,558 words

I've spent enough years reading consent orders to know the pattern by now: examiners don't ask if you filed paperwork on time anymore. They ask whether the system flagging a transaction, or talking to a customer, can be explained after the fact, in plain language, to someone who wasn't in the room when it happened. That's the real story in the 2024 to 2025 enforcement record, and it's worth walking through slowly, because the details tell you more than the headline numbers do.

Guideposts Solutions counted 44 enforcement actions from the start of 2025 through May. That's a fast pace, and it tells you something specific: wherever new technology runs up against old compliance obligations, examiners are watching harder than they used to. The SEC got there first. Back in 2023, it ran exams on roughly 30 registered investment advisers focused specifically on AI disclosures and governance, the first time AI conduct became its own line item on an exam rather than an afterthought. What they found were firms with no real AI policy at all, and in a few cases, outright misstatements about how the AI was actually being used. That pattern didn't stay in securities. It walked straight into banking supervision, and it's still walking.

Why should anyone outside the named firms care? Because consent orders put the examiner's reasoning on paper in a way a routine exam letter never does. Read enough of them back to back and a rulebook starts to form, one no single agency sat down and wrote, but one the whole industry now has to answer to anyway.

The penalty benchmarks that define the cost of getting this wrong

Start with the number everyone remembers: FinCEN's $1.3 billion penalty against TD Bank in October 2024, the largest in the agency's history. The order described "chronic failures" in the bank's anti-money-laundering program. Examiners weren't grading one bad quarter. They were checking whether the governance held up over years, and it didn't hold.

The Federal Reserve's $43 million penalty against Silvergate Bank zeroed in on deficient transaction monitoring, which happens to be the exact control banks are now racing to hand off to AI. Worth sitting with that for a second. Regulatory actions against larger institutions over unsafe BSA/AML and sanctions controls showed the same expectations apply whether you're a regional bank or one of the largest in the country. And the FDIC's $20.4 million penalty against a Kansas bank tied weak AML and CFT programs to $27 billion in annual wire flows. Volume alone draws scrutiny, no matter how simple the rest of the operation looks on paper.

Here's the part that surprised me: total public enforcement actions actually fell, from 17 in 2023 to 12 in 2024, with about $48.8 million in combined penalties that year. Fewer cases, a smaller combined total, and yet the single largest penalty in FinCEN's history landed right in the middle of that same stretch. Lower volume doesn't mean softer expectations. It means regulators are picking their targets more carefully and hitting harder once they find what they're looking for.

Diagram: Enforcement Volume Fell While the Largest Penalty Ever Landed. Visualizes: Visualize the paradox at the heart of 2024 FinCEN enforcement: total public enforcement actions dropped from 17 in 2023 to 12 in 2024, and combined penalties fell…

The specific deficiencies examiners documented — and what they imply about AI oversight

K&L Gates reviewed 2024 BSA/AML enforcement actions and found independent testing failures showing up as a prominent deficiency. In at least one action, examiners found the testing scope simply didn't match the institution's real risk profile.

Look closer at what "testing failure" meant in practice. The testing confirmed a control existed, but it didn't confirm the control was doing anything. That distinction sounds minor until you apply it to AI. A transaction monitoring model with clean documentation and a tidy architecture diagram tells an examiner nothing about whether it catches what it's supposed to catch once it's live. Documentation of existence isn't proof of performance, and examiners are saying that out loud now, not just implying it.

Governance gaps showed up too, mostly as AML officer roles that existed on an org chart but carried no real authority behind them. Somebody's name sat in the box, while nobody was actually on the hook for what the automated systems produced. SAR and CIP timeliness issues aren't new by themselves, but they carry new weight once AI enters the workflow. Automated handoffs bring their own delays and their own failure points, and an examiner isn't going to treat a lag caused by an AI pipeline any more gently than one caused by a slow analyst. A control failure is a control failure, whatever produced it.

The SEC's 2023 findings on AI disclosure, misrepresenting how AI gets used or not documenting it at all, are turning up again in banking exam scope. Voice AI is increasingly running collections calls, delivering disclosures, and handling escalations in live customer conversations. So examiners are asking something direct: which conversations did the AI actually handle, and can a supervisor reconstruct and defend what it said afterward?

Put these findings side by side and one thread runs through all of them. Examiners are taking standards that have existed for years and pointing them at automated systems, and finding that the documentation, the testing, and the human accountability weren't built to keep pace with what the technology was already doing.

What the April 2026 interagency model risk guidance actually changed — and what it left unresolved

SR 26-2 and OCC Bulletin 2026-13, both issued April 17, 2026, replace SR 11-7, the model risk framework that governed bank AI review for over a decade. The headline change: generative and agentic AI are explicitly carved out of the revised guidance's formal scope. The agencies called these technologies "novel and rapidly evolving" and said a separate request for information is coming.

So does sitting outside the scope mean sitting outside accountability? No, not according to the guidance. The guidance says existing risk management principles still apply to tools sitting outside its formal boundaries. The practical read is plain: governance practices should still shape what controls a bank applies, even to AI the guidance doesn't formally cover.

What that means day to day: the evaluation burden sits with the institution now, not with a checklist a regulator handed down. An examiner will ask what governance a bank applied to its agentic AI, and "the guidance didn't require it" is not going to land well as an answer.

FDIC, Federal Reserve, OCC, and NCUA officials told GAO that AI already gets reviewed under existing safety and soundness, IT, and compliance exam frameworks, whenever risk-based planning flags it as relevant. The hook to examine AI already exists; it doesn't need a new rule to switch on. The CFPB confirmed something similar: it examines AI through product and practice reviews rather than a standalone AI lens, which means a customer-facing AI agent gets measured against fair lending, disclosure, and UDAAP standards that were written with human employees in mind, not machines. And the OCC's May 2026 Semiannual Risk Perspective treated AI as both a cybersecurity threat and an operational risk, which stretches the exam surface well past model risk alone.

Where NCUA's oversight framework falls short and what that means for credit unions deploying AI

GAO's February 2025 report found NCUA's model risk guidance too thin, both in scope and in detail, for staff or institutions to actually understand how AI model risk should get managed. That's bad enough on its own. The bigger problem sits underneath it: NCUA doesn't have the statutory authority to examine technology service providers directly, even as credit unions lean harder every year on third-party vendors to run their AI.

GAO had already recommended Congress grant that authority, and as of February 2025, Congress hadn't moved on it. Think through what that gap actually means on the ground. A credit union's AI vendor could have a governance failure baked into the platform itself, and NCUA has no direct way to examine that vendor. The credit union still carries the regulatory exposure, even though the failure started somewhere NCUA can't reach.

To be fair, NCUA hasn't sat on its hands. It published an AI compliance plan and brought on dedicated AI officers for 2025 and 2026, signaling that the bar is guardrails, accuracy, transparency, and auditability. But a thinner guidance framework doesn't translate into thinner examiner expectations; it just means credit unions have to build their own rigor instead of leaning on a detailed regulatory blueprint that, for banks, at least partly already exists. Credit unions face the same staffing pressure, the same competition from digital-first players, and the same member expectations pushing AI adoption forward as banks do. They're just doing it with less scaffolding underneath them.

How rapidly banks are moving into agentic AI relative to how slowly governance is following

Wolters Kluwer's Q1 2026 survey of 148 financial institutions found roughly 31.8% had already put AI or machine learning into production. Only 12.2% described their AI strategy as well-defined and properly resourced. Sit with that gap for a second: nearly a third running these systems live, barely one in eight with a strategy that actually matches the deployment. That gap is exactly where I'd expect the next round of enforcement findings to come from.

The same survey found 44% of finance teams expect to use agentic AI in 2026, a jump steep enough that the researchers themselves flagged it as coming "potentially at the expense of clear strategy and AI governance." North America already holds the largest share of the global agentic AI market for banks and credit unions, at 42.3% in 2025, worth roughly $2.03 billion in revenue. The technology is scaling fast, and governance isn't keeping the same pace, not even close.

One might argue the efficiency case justifies moving quickly anyway. Global compliance costs for banks now run north of $270 billion a year, and agentic AI genuinely can cut into that number. But efficiency bought through automation nobody can audit afterward is exactly the liability profile these enforcement actions keep punishing; the savings and the exposure come from the same place. Add a vendor market where rival platforms have rolled out "industry-first" agentic banking claims within weeks of each other across 2025 and 2026, and the real task facing any institution is figuring out what the AI actually does in production, and what controls are actually built in rather than bolted on afterward.

Diagram: Deployed vs. Ready: The AI Governance Gap in Financial Institutions. Visualizes: Show the stark deployment-versus-readiness gap from Wolters Kluwer's Q1 2026 survey of 148 financial institutions: 31.8% had already put AI or machine…

The four governance structures that enforcement findings reveal as non-negotiable

Independent testing scaled to real risk comes first. The K&L Gates review found testing flagged as deficient whenever it confirmed a control existed without confirming it worked. Applied to AI, that means testing has to show the model produces the outputs it was built to produce, at the volume and risk level the bank is actually running, not the volume it ran during the pilot. Scope, frequency, and method need to be written down in a way an examiner can follow without someone standing over their shoulder explaining it.

Then there's the three lines of defense, applied to the AI agent itself rather than bolted onto an org chart drawn up before the agent existed. Line one is the controls built directly into the agent: transaction limits, escalation triggers, confidence thresholds, and full logs of every action taken. Line two is an independent risk function reviewing the model's performance and escalated cases, with real authority to shut the thing down if it needs shutting down. Line three is internal audit, checking that line one's controls actually work and that line two is independent in practice, not just independent on the chart. The failure pattern regulators keep circling back to: all three lines sitting under one technology team, with no real separation anywhere. That's the same structural problem behind the "qualified AML officer" deficiency from the 2024 enforcement actions, just wearing a different hat.

Audit trails need to get built in at deployment, not bolted on after the fact. Examiners want to know which conversations the AI handled, what it said during a disclosure or a collections call, and whether a supervisor can reconstruct and defend those decisions later. Emerging documentation standards — covering who deployed the system, what data went in, what decision came out, and what human oversight actually happened — are becoming a working benchmark even for institutions with no formal obligation to follow them. SR 26-2 and OCC Bulletin 2026-13 have effectively turned this into a procurement requirement. Check a vendor's log completeness before you sign the contract, not after an examiner shows up asking questions you can't answer.

And vendor governance has to cover the AI itself, not just the vendor relationship in the abstract. The joint OCC, Federal Reserve, and FDIC third-party risk guidance from 2023 applies to AI vendors the same way it applies to any other vendor: planning, diligence, ongoing monitoring, and accountability don't become optional just because the AI happens to run on someone else's servers. NCUA's inability to examine service providers directly makes this even more urgent for credit unions, since the institution carries the exposure no matter who's actually running the model. At minimum, vendor evaluation should cover documented access controls and authentication, granular logs of every action the AI takes, and real evidence that governance standards are met in practice, not a polished sales deck with a logo on it. Institutions building agentic AI on top of existing banking rails, with configurable controls and full audit trails from day one, start from a stronger position than those bolting general-purpose AI onto core systems after the fact.

What institutions can do now before the next examination cycle closes the window

The revised interagency guidance leaves agentic AI in a kind of holding pattern; the RFI the agencies pointed to hasn't been issued yet. Institutions waiting around for it are still getting examined under existing principles in the meantime, whether they've prepared for that or not.

Start with an inventory. Map every AI system currently in production to an existing exam framework: safety and soundness, IT, compliance, or consumer protection. That's how GAO's February 2025 report says examiners will actually scope their review, so use that lens now, before an exam forces the question on you.

Next, pull the specific deficiency language from the 2024 enforcement actions and stress-test your own independent testing documentation against it directly. Does the evidence show the control is functioning, or does it only show the control exists? Those are different questions with different answers.

Then audit the audit trail itself. Do the logs actually capture what an examiner would need to reconstruct a decision the AI made, whether that's a voice call, a monitoring alert, or an automated transaction? If the honest answer involves a shrug or a "we'd have to go check," that's the gap to close first, before anything else on this list.

Last, go back through every AI vendor contract against the 2023 joint TPRM guidance. Diligence, ongoing monitoring, and accountability provisions need to be spelled out in writing, not assumed because the vendor seems reputable or because everyone else in the industry uses them too.

Institutions that turn the 2024 and 2025 enforcement record into real governance changes this year are cutting their penalty exposure directly. But there's a second payoff that matters just as much: they're building the kind of internal credibility that lets AI deployment keep expanding with the regulator's confidence behind it, instead of running one step ahead of the next consent order.

Sources

  1. anaptyss.com
  2. guidepostsolutions.com

More in Features