Rail Governance

Liquidity Risk Examination Expectations for Banks Running Automated Payment Agents

Regulators expect banks to govern automated payment agents despite unfinished rules on AI.

Senior Writer · · 11 min read
Cover illustration for “Liquidity Risk Examination Expectations for Banks Running Automated Payment Agents”
Examiner Expectations · October 3, 2026 · 11 min read · 2,500 words

Picture the old treasury workflow. A treasury professional logs into a portal, checks the day's cash position, decides whether to move money, and then moves it. That gap between decision and action, usually minutes, sometimes longer, was never written down as a control anywhere. But it worked like one. It gave a human a last chance to catch something wrong before money left the building. Automated payment agents close that gap. Agents built on instant payment rails, paired with programmable money and preset conditions, can move funds around the clock without a person in the loop. The friction that informal controls quietly depended on is gone, and nothing has formally replaced it.

The BCG Global Payments Report 2026, published September 23, 2026, documents this shift in plain terms: AI agents are already routing payments in real time based on cost, speed, and straight-through-processing criteria, managing currency exposure and cash pooling inside a company's defined policies. Cross-border flows show the same pattern. FX conversion, routing choice, and settlement timing are increasingly decided by machines within seconds, not by a person reading a screen. Once decisions move at that speed, risk teams, compliance teams, and liquidity managers have no choice but to operate at that speed too, or fall behind the thing they're supposed to be watching.

The IMF frames the deeper issue well in an April 2026 paper: payment systems are built on deterministic rules, the same input always produces the same output, in a sequence that regulators and operators can trace step by step. Agentic AI doesn't work that way. It behaves adaptively and probabilistically, weighing options and producing outputs that can vary even when conditions look similar. Those are two fundamentally different ways of making decisions, and the IMF is direct that they don't fit together without extra safeguards built in on purpose. That mismatch, not a vague sense that "AI is different," is the structural reason intraday liquidity monitoring built for human-paced flows cannot simply be pointed at agent-driven flows and expected to work. The rest of this piece works through what that mismatch means in practice, starting with how it plays out when agents aren't acting alone.

Correlated, cascading outflows from agent-driven payments that stress contingency funding plans

A single misbehaving agent is a manageable problem. A bank can build a control for that, cap it, log it, shut it down. What happens when many agents move the same way at the same time, for reasons that make sense to each of them individually, is harder to control. S&P Global's 2025 research flags exactly this kind of exposure: automation now links trading, credit, and compliance systems across institutions through continuous data exchange, and that connectivity can amplify volatility when agents respond to the same shock in correlated ways. No single agent has to misbehave for the system to wobble. They just all have to react the same way, at the same moment, to the same signal.

Chaining makes the problem move faster than a person can react to it. Agents increasingly call other agents and chain tools together to complete a task. A single bad input, a mispriced trade, a duplicated payment instruction, a misrouted transfer, doesn't stay contained to one workflow. It can multiply across systems before anyone sees the first alert, let alone acts on it.

The Bank for International Settlements tested something close to this directly in 2025, simulating a generative AI agent managing cash in a real-time gross settlement system. Taken alone, the agent behaved the way a careful human cash manager would: holding buffers, prioritizing payments sensibly, managing liquidity constraints without unnecessary cost or delay. But the BIS researchers found that individual prudence doesn't guarantee collective safety. Once multiple participants in the same system are coordinating, or failing to coordinate, their combined behavior affects gridlock risk and liquidity externalities in ways that no single agent's good behavior can offset. A room full of cautious agents can still produce an outcome nobody designed and nobody wants.

This is the scenario most contingency funding plans don't test for. Standard CFP stress scenarios model depositor runs and market dislocations happening at human speed, where withdrawal requests arrive over hours or days and a bank has time to activate funding lines, call counterparties, and adjust. Agent-driven, correlated outflows don't follow that timeline. They happen in seconds, and they can come from agents running the same underlying model making the same decision near-simultaneously, inside one institution or across several. That's a gap in the template, not a hypothetical one, and it's the gap the rest of this piece works toward closing.

Where guidance gaps leave banks exposed

Supervisors have moved, but they haven't moved all the way to a finished rulebook, and that gap matters more than it might first appear. On April 17, 2026, the Federal Reserve, the OCC, and the FDIC issued revised interagency model risk management guidance, designated SR 26-2 by the Federal Reserve and Bulletin 2026-13 by the OCC. It replaced the older SR 11-7 and SR 21-8 guidance that banks had relied on for over a decade. The agencies said that generative and agentic AI are novel and still evolving fast enough that they fall outside the guidance's formal scope. The agencies signaled they plan to issue a request for information specifically on banks' use of AI, which tells examiners and banks alike that more is coming, just not yet.

What the agencies didn't do is leave a vacuum. They were explicit that the core principles, materiality, ongoing monitoring, and effective challenge, still apply to tools that sit outside the new guidance's formal boundaries. That means a vendor marketing an agentic payment platform as "SR 26-2 compliant" is describing a status the guidance itself doesn't actually confer, since agentic AI isn't inside its scope to begin with.

State supervisors moved next. On September 16, 2026, the Conference of State Bank Supervisors released an AI supervisory framework built for examiners reviewing state-chartered banks and state-licensed nonbank financial institutions. It hands examiners specific questions to ask, specific documents to request, and criteria for deciding when an AI system needs closer scrutiny. The framework's underlying logic is that the risk isn't that a bank uses AI. The bank can't prove, when asked, that the AI is governed, controlled, and compliant.

The IMF's April 2026 paper makes a related point about the regulatory picture overall: it's deliberately unfinished right now, and that incompleteness functions as its own risk signal inside a supervisory relationship. An examiner looking at a gap in formal rules doesn't read it as permission. The Financial Stability Board's 2026 consultation on responsible AI adoption pushes in the same direction internationally, proposing sound practices for AI governance and lifecycle management across an entire organization, and asking directly whether those practices can stretch to cover emerging forms like generative and agentic AI.

Credit unions sit inside a similar posture. The NCUA's 2026-2030 Strategic Plan, under Strategic Objective 2.1, commits to fostering an environment where federally insured credit unions can adopt financial technology, digital assets, and other innovations responsibly. That's a supportive stance, but the word "responsibly" carries weight: it implies scrutiny of how that adoption is governed.

Put together, examiners are applying frameworks built for a human-paced banking world to systems now running at machine speed, and they're doing it before agent-specific rules exist to guide them. A bank that can't document how its agents are governed will face findings under the liquidity and model risk principles already on the books, with or without a final agentic AI rule to point to.

Intraday liquidity monitoring and the invisibility of agent reasoning

Watching the outcome of a payment isn't the same as understanding how an agent got there, and that distinction is where a lot of monitoring programs quietly fail. Two separate ideas matter here. Explainability answers why an action happened. Reproducibility answers whether anyone could rebuild the same decision process later, under the same documented conditions, and get the same result. Agentic AI weakens both at once because the decision is spread across multiple models, multiple tool calls, and handoffs between agents, each step adding a link a reviewer has to trace.

A monitoring system that only sees the final transaction, a debit posting, a transfer confirmation, can't tell an examiner much that actually matters. It can't say whether the agent stayed inside its configured limits. It can't say whether the agent deliberately delayed a lower-priority payment to protect a liquidity buffer, which would be good behavior, or whether it was reacting to a prompt injection or a model error, which would be a serious problem wearing the same outward appearance. Both look identical from the outside if all a bank is logging is the settled payment.

The leaked episode involving Nvidia and Bank of America illustrates how this gap appears in practice. Even a major institution deploying advanced AI can run into a simpler, more basic problem: staff lacking the operational fluency, particularly the MLOps skills, needed to actually run and manage what's been deployed. That means a monitoring system and an agent can both be technically running, fully switched on, generating logs, while the institution has no real working visibility into what either one is actually doing. The tooling existing isn't the same as the tooling being understood.

That's the gap the next section tries to close: not by adding more dashboards, but by designing the controls into the agent's workflow from the start, so the reasoning behind an action is captured at the moment it happens rather than reconstructed afterward from whatever logs happen to survive.

Configurable controls, real-time visibility, and complete audit trails for examiner expectations

Retrofitting governance onto an agent that's already deployed rarely works, because by the time a problem occurs, the reasoning behind the agent's decision is often gone. The fix has to be built into the system before it goes live, with three things in place from day one: human checkpoints at key decision points, settings a bank can actually adjust, and a complete record from instruction to outcome.

High-value or unusual outflows need a human-in-the-loop checkpoint before execution, not after. If the agent pauses and routes an anomalous payment to a reviewer before money moves, that's a control. If a human only reviews the payment after it's already settled, that's a report, and examiners draw a real distinction between the two. Payment velocity limits, counterparty concentration caps, and intraday liquidity floor thresholds all need to be settable by the institution itself, with every change to those settings logged, because an examiner's first question is usually whether limits exist, and the second is whether they were actually enforced at the moment the payment went out, not just written down in a policy binder somewhere.

Every tool call, every handoff between agents, and every reasoning step needs to be logged in a form that lets an examiner walk the full chain backward, from the original instruction to the final settled transaction. Logging only the end result doesn't hold up under that kind of review. Vendor oversight has to stretch to cover the agents themselves as well as the platforms running them. If a core provider's agent, Fiserv's agentOS or FIS's Financial Crimes AI Agent are two examples already live in the market, is initiating or shaping a payment decision, the bank's vendor management program needs to cover that agent's governance directly, not just confirm the platform stays up and running.

Hua et al., writing in April 2026, offer one useful model for how this kind of accountability can be built structurally rather than left to policy language. Their Agentic Risk Standard turns the uncertainty around an agent's outcome into explicit, enforceable settlement rules: service fees held in escrow until conditions are met, collateral posted before the full principal changes hands. The specific mechanism is theirs, but the underlying idea, that compensation and outcomes should be tied together by predefined, contractually enforceable rules, lines up closely with the kind of documented control environment examiners are already looking for.

None of this matters if the records can't be produced on request. The CSBS supervisory framework gives examiners a specific list of documents to ask for, and a bank that can't produce them will face findings regardless of how well its agents actually performed in practice. SOC 2 certification is worth having, but it covers data security and operational controls generally. It isn't built for agent-specific governance questions, and examiners will treat it as a baseline, not as proof that the agent itself is well governed.

One practical choice shapes how hard all of this is to deliver: whether an agent runs on a bank's existing, already-audited banking rails, or on separate infrastructure built just for it. Keeping agent activity inside systems that are already subject to established controls and existing audit processes cuts down the documentation burden considerably, for both model validation and liquidity reporting, compared to standing up a parallel system examiners have never looked at before.

Making contingency funding plans and stress tests credible against agent behavior the standard templates do not cover

The governance work described above only closes part of the loop. Contingency funding plans and stress tests are where a bank proves, on paper and under questioning, that it has actually thought through what agent-driven liquidity risk looks like under pressure, and right now, most standard templates haven't caught up. Existing CFP scenarios model depositor runs and market dislocations unfolding at a human pace, over hours or days, giving a bank time to pull funding levers in sequence. An agent can look careful and well-behaved in isolation, and the same agent's behavior can still contribute to gridlock once it's interacting with other participants making similar choices at the same time, which is the behavior the BIS simulation already demonstrated. A credible CFP has to test for that kind of correlated, synchronized outflow directly, checking whether intraday buffers actually hold up when multiple automated systems move together rather than one at a time.

Stress testing needs to go further than correlated outflows alone. A bank should be able to show what happens when an error, a model mistake or a prompt injection, moves through a chained agent workflow faster than a human reviewer can catch it, and how long it actually takes to recover once the stage-gate controls described earlier do their job and stop execution. The same discipline applies to vendor concentration: if a single model provider powering several banks' agents goes offline, or pushes out a flawed update, a bank needs a tested answer for how it reverts to manual controls and how long that transition realistically takes under pressure, not just a policy statement that manual fallback exists. None of these scenarios require exotic assumptions. They follow directly from what agents are already doing in production, at a scale that's only going to grow, which makes updating the CFP and stress test templates less a matter of future-proofing and more a matter of catching the paperwork up to where the risk already sits.

Sources

  1. Global Payments Report 2026: Transaction Banks Must Architect for an Agentic Future
  2. How Agentic AI Will Reshape Payments in: IMF Notes Volume 2026 Issue 004 (2026)
  3. AI agents for cash management in payment systems
  4. Quantifying Trust: Financial Risk Management for Trustworthy AI Agents
  5. OCC Issues Updated Model Risk Management Guidance
  6. August 2026 Global Regulatory Brief: Bank capital, liquidity risk management and regulatory modernization
  7. Modeling Trust and Liquidity Under Payment System Stress: A Multi-Agent Approach
  8. AI Agents in Payments: Applications, Risks and Regulations

More in Examiner Expectations