Rail Governance

BSA Examination Exposure from AI-Generated Transaction Narratives

Examiners are finding control gaps in AI-drafted suspicious activity reports filed by banks.

Staff Writer · · 9 min read
Cover illustration for “BSA Examination Exposure from AI-Generated Transaction Narratives”
Examiner Expectations · September 13, 2026 · 9 min read · 2,121 words

U.S. financial institutions filed 4.7 million suspicious activity reports in fiscal year 2024. That works out to roughly 12,870 a day, and every one of them needs a narrative spelling out who did what, when, where, why, and how. AI now writes a growing share of those narratives, and the technology has opened a gap between what examiners expect and what institutions have actually built to control it. That gap, not whether AI belongs in SAR drafting at all, is what this piece is about.

How AI entered SAR drafting and what the first generation got wrong

SAR Form 111 isn't a form you fill in loosely. The FFIEC BSA/AML Examination Manual treats completeness and thoroughness as explicit requirements, not suggestions. A narrative has to name the typology, walk through the method of operation, and cite the source systems behind every claim. In October 2025, FinCEN's FAQs sharpened that further, shifting the regulator's public framing from quantity of filings toward quality, on the theory that law enforcement needs filings it can actually act on.

That's a task shape AI happens to fit. A reasoning model, fed the right case data, can draft a narrative, propose a typology classification, and cite its sources, while a human reviews and files. A good SAR narrative already follows a repeatable pattern (who, what, when, where, why, how, typology, citation), so the fit isn't an accident.

First-generation tools built for this job stumbled on one specific defect: they hallucinated. Counterparty names that didn't exist in the case file. Transaction details invented to fill a gap in the narrative. When a model can't trace a claim back to its alert source, the problem it creates is legal. Wolters Kluwer's compliance guidance names the risk: hallucinated facts, omitted details, bias baked into poor training data, reduced analyst oversight. Any one of those can weaken an investigation, produce an inaccurate filing, or draw regulatory scrutiny.

By 2026, a clearer picture had emerged of what actually separates a defensible tool from a defective one. Every claim is grounded in retrieval from the case file, each assertion has a citation behind it, and human review before filing is mandatory. That's a dead end rather than an upgrade path. It's the floor. A tool that skips any one of those three isn't a lesser version of a good system; it's a different, riskier category of tool.

The three AI-specific vulnerabilities examiners can identify in a filed SAR

Start with the obvious one. A model generating narrative text without retrieval grounding can produce a counterparty name, an account number, or a transaction amount that doesn't appear anywhere in the underlying alert. Examiners comparing the filed narrative against the investigation record will find the mismatch, and there's no good answer for a fact the institution can't source.

How hard is full traceability to pull off, even with a purpose-built system? A research framework called SCAFDS (arxiv 2605.18913) tested an attribution-conditioned SAR generation system and found traceable assertions in only 60% of cases. That means 40% of the claims in an engineered system, one built from the ground up for attribution, had no per-claim link back to a specific numerical output. If that's the miss rate on a system designed for this exact job, a general-purpose drafting tool bolted onto an existing workflow deserves real scrutiny, not the benefit of the doubt.

Second: uncitable attributions. AI fraud detection models throw off SHAP feature attributions and ensemble probability scores, and compliance staff sometimes fold those numbers straight into the SAR narrative as justification for the filing. FinCEN hasn't issued guidance on how a model output like a SHAP score should be used in a narrative, or what an examiner is supposed to make of it once it's there. So an institution that cites a model confidence score but can't produce the model's documentation, its validation record, or an audit trail of its outputs has written a claim its own compliance officer can't explain when asked.

Third: typology gaps that only surface in aggregate. Examiners are starting to check whether institutions can spot AI-enabled fraud patterns in their own SAR narratives, and FinCEN has signaled more interest in AI fraud risk disclosure, with guidance expected in 2026 and 2027 likely to address AI-enabled identity fraud typologies directly. A drafting tool trained on a dataset that predates current fraud patterns won't just miss cases one at a time. It'll misclassify a pattern across the whole filing population, and that pattern is exactly what an examiner looking at the full set is trained to spot.

Where SR 26-2 and FinCEN's April 2026 proposed rule leave the governance gap

On April 17, 2026, the Federal Reserve, the OCC, and the FDIC issued SR 26-2, the first full rewrite of model risk management guidance in fifteen years. It replaces SR 11-7 and supersedes the 2021 Interagency Statement on Model Risk Management for BSA/AML systems.

Here's the catch: SR 26-2 explicitly excludes generative and agentic AI models from its scope, calling the technology "novel and rapidly evolving." So the transaction monitoring models, sanctions screening systems, and CDD tools that feed a SAR get inventoried and tiered under the new rule. The AI layer that takes those outputs and drafts the actual narrative sits outside it.

That exclusion is not a green light, whatever it might look like on first read. Analysis of SR 26-2 suggests examiners are already asking every bank, regardless of size, how it governs the AI systems the rule doesn't cover. The rule's silence on generative AI has become the examiner's opening question, not a reason to skip it.

FinCEN's proposed rule from the same month pushes in a similar direction. It actively encourages the responsible use of machine learning and generative AI in BSA workflows, while making clear that responsible adoption is the operative standard. But "responsibly" is doing real work in that sentence. The institution and its compliance officer stay accountable, a human makes the filing decision, and the whole process has to be governed and documented on paper, not just in practice.

So where does that leave things in mid-2026? The framework is being rebuilt, not tightened along one straight line, and generative AI in BSA workflows sits in the gap between an old rule that never anticipated it and a new rule that carves it out on purpose. Reading that gap as permission is the wrong move. Meanwhile, the OCC's November 2025 update to community bank BSA/AML examination procedures leaned toward more examiner discretion and risk-based scoping, which tells you the examination side of the house is already moving faster than the rulebook.

Diagram: The Governance Gap: What SR 26-2 Covers — and What It Doesn't. Visualizes: Illustrate the regulatory blind spot created by SR 26-2 (issued April 17, 2026 by the Federal Reserve, OCC, and FDIC).

What accountability actually means when an AI system drafts the narrative

The rule here is simple to state and easy to underestimate: the institution and its BSA officer are accountable for every filed SAR, no matter what produced the draft. That accountability doesn't move to a vendor, and it doesn't move to the model, no matter how the contract with the AI vendor is worded.

Jo Brown puts it directly: regulators have made clear, through guidance, through examinations, through enforcement actions, that automation doesn't reduce accountability. If anything, oversight expectations go up as the AI gets more capable, not down.

So what does human review actually require? Not a signature at the bottom of a draft. Analysts have to check and confirm AI-generated outputs, especially at high-risk decision points where errors carry the greatest regulatory or legal consequence. A rubber stamp doesn't clear that bar, and an examiner who spots a run of unchanged AI drafts across a sample of filings is going to ask what the review actually consisted of.

"Agent drafts, human attests" looks like the pattern regulators are comfortable with long-term. Whether it functions as a real control or just a formality, though, comes down entirely to what the attestation involves.

The industry baseline right now isn't close to where it needs to be. A 2025 Infosys study found only 2% of companies had adequate AI guardrails in place, while 95% had experienced at least one AI incident, 77% of which led to financial losses and 55% to reputational damage. IBM's 2025 Cost of a Data Breach report found 63% of organizations had no AI governance policy at all, and among organizations that suffered an AI-related security incident without proper access controls, 97% found shadow AI involved. That last number matters for BSA teams specifically: tools adopted outside the formal model governance process are exactly the shadow AI risk IBM is describing.

The four governance controls that convert AI-assisted drafting from a liability into a documented program

Model governance with defined ownership. Any AI system touching alert adjudication, risk scoring, or narrative drafting belongs inside the institution's model risk management framework: named ownership, a validation protocol, ongoing monitoring, data quality testing. SR 26-2 covers traditional models while explicitly excluding generative and agentic AI from its scope. The drafting layer sitting above the scoring models should get equivalent treatment even though the rule doesn't demand it, because the examiner will ask anyway. Treasury's AI Risk Management Framework, built around four functions (Govern, Map, Measure, Manage) with roughly 230 control objectives scaled to an institution's stage of AI adoption, gives institutions without their own framework something examiner-credible to point to.

Per-claim grounding with traceable citations. Every assertion in a machine-generated narrative needs a source: a specific document, transaction record, or system output in the case file. This is the core requirement, the actual architectural difference between the hallucinating first-generation tools and the grounded-retrieval systems that replaced them, rather than a nice-to-have layered on top of a good tool. An examiner who pulls the case file and can't find where a narrative claim came from has found a defect, full stop. Institutions need to decide, in advance, what counts as a citable source, what happens when a model can't ground a claim, and how the human reviewer checks citations before signing off.

Human attestation that's actually substantive. The reviewing analyst has to verify the narrative against the case file, not just confirm a draft exists. And the review needs a paper trail: what got checked, what got changed, and why. An attestation with no documentation behind it looks, from an examiner's chair, exactly like no review happened at all. High-risk points, structuring patterns, OFAC hits, typologies the institution hasn't seen before, need an explicit confirmation step, never an assumed one.

Audit trails built to survive examination. Model version, grounding sources, the analyst's review record, whatever changed between the draft and the filed narrative: all of it needs to be kept and produced on request. Examiners already compare SAR filings against an institution's policies, risk assessments, and case records, and the audit trail is what ties the filed document back to that context. Reported data suggests the average enterprise bank's AI governance budget grew 62% in 2025. Institutions that haven't matched that pace in audit infrastructure are likely carrying exposure they haven't measured yet.

What examiners will look for as AI SAR drafting becomes standard practice

The OCC's November 2025 community bank procedures already touch on AI use in SAR narrative generation. Sit with that for a second: the examination apparatus is asking AI-specific questions right now, not preparing to someday.

As AI drafting spreads, the examiner's question shifts. It stops being "did you use AI" and becomes "how did you govern it": model documentation, grounding standards, the substance of the attestation, whether the audit trail is complete. FinCEN's expected 2026-2027 guidance on AI-enabled identity fraud typologies will likely set a new bar for classification accuracy, and institutions whose tools are still classifying off an outdated training set will have a specific, documentable problem on their hands, not an abstract one.

Wolters Kluwer's guidance offers a fair place to land: in a period where the rulebook is being rewritten, a well-documented, risk-based program is a real advantage, not just a compliance cost. Institutions building that governance now, ahead of the formal rule catching up, will face less friction at examination than the ones waiting for a final rule to tell them what to do.

The institutions in the strongest position aren't treating AI SAR drafting as a productivity add-on bolted onto an existing workflow. They're treating it as a model deployment, full stop: same ownership requirements, same validation discipline, same monitoring, same audit trail as any other system whose output a regulator is going to read line by line. Systems built to run on existing banking rails while keeping a complete audit trail at the transaction and output level are simply better matched to that posture. Configurable controls, complete records, a human at every real decision point: those aren't extra features when the examiner is the intended reader. They're the whole point.

Sources

  1. Five developments every compliance leader needs to know | Wolters Kluwer
  2. arxiv.org
  3. arxiv.org
  4. The Effective Use of AI for SARs | ACAMS
  5. cimcon.com
  6. sullcrom.com
  7. federalreserve.gov

More in Examiner Expectations