Rail Governance

Third-Party Risk Management for AI Vendors Executing Banking Transactions

Banks must clarify who owns failures when AI agents execute transactions without human approval.

Features Editor · · 12 min read
Cover illustration for “Third-Party Risk Management for AI Vendors Executing Banking Transactions”
Examiner Expectations · September 9, 2026 · 12 min read · 2,755 words

Money moving on its own changes what third-party risk management has to check for. Standard vendor risk frameworks were built to review advisors: people or systems that suggest a move and wait for a human to say yes. An AI agent that executes the transaction itself doesn't wait, and most bank TPRM programs still don't have a clean answer for what happens when that agent gets it wrong.

Here's the accountability question that keeps surfacing, and most vendor risk programs still can't answer it cleanly: if the agent executes a transaction and it's wrong, which control failed, and who owns that failure? Legal analysts flagged a version of this problem in late 2025, warning that third-party AI agents acting on behalf of customers carry real legal exposure when contracts, human approval gates, and liability lines haven't been spelled out ahead of time. The same gap opens up when the agent acts on behalf of the bank itself, not the customer. Third-party risk management for transaction-executing AI has to start with execution accountability. Treating vendor due diligence as a box to check after the deal is signed gets the sequence backwards, and it's the mistake most banks are currently making.

How quickly transaction-executing AI is moving from pilot into production

This isn't a someday problem. Roughly 42% of financial services firms are currently using or actively assessing agentic AI, according to the State of AI in Financial Services 2026 Trends report. That sits on top of much broader AI adoption across the sector (81% of firms use AI in some form, and a significant share report meaningful deployment). Agentic transaction execution is the next rung up from general-purpose AI tools, and it's a taller rung than most banks are treating it as.

The IMF's April 2026 outlook described agentic AI as capable of running a full cross-border payment chain on its own: initiation, routing, compliance checks, settlement monitoring, all handled by the same system. The IMF was careful to call this early, experimental territory, not settled practice. Related analysis in 2025 noted that AI agents are being applied to liquidity management and payment prioritization inside real-time gross settlement systems. That's not a chatbot answering customer questions. That's software deciding, in real time, what order money moves in.

Oversight hasn't caught up, and the gap is wide enough to matter. IBM found in late 2024 that only 8% of banks build AI in a genuinely strategic, enterprise-wide way, while 78% stay stuck in tactical, one-off deployments. McKinsey's 2026 AI Trust Maturity Survey put a number on the governance gap directly: an average responsible-AI maturity score of 2.3, up from 2.0 the year before, with only about a third of organizations scoring a 3 or higher across strategy, governance, and agentic AI oversight. Execution capability is sprinting. Oversight is still lacing up its shoes. Closing that gap is the actual job of a transaction-focused TPRM framework, not a side benefit of one.

Where existing TPRM guidance reaches its limits with agentic vendors

The regulatory baseline still applies, and nobody gets to skip it. OCC and FFIEC third-party risk management guidance from 2023, the model risk management standard now known as SR 26-2 (which replaced SR 11-7 in April 2026), and OCC Bulletin 2025-26 for community banks all sit underneath any agentic AI deployment. Comptroller Jonathan Gould has said plainly that AI should be governed by technology-neutral, risk-based principles, and that the existing framework holds. But he's also been clear that "proportionate does not mean absent." A smaller scope doesn't buy a bank a pass on oversight.

Here's where the fit gets awkward, and it's the part most compliance teams gloss over: SR 11-7 was written for static, batch-run models, the kind you validate once and re-check on a fixed schedule. An agent that keeps learning after deployment doesn't hold still long enough for that kind of validation to mean anything. If the model behind a transaction-executing agent looks different in March than it did in January, how much is that January validation actually worth by the time March rolls around?

Standard vendor due diligence asks whether the vendor has strong controls. Agentic due diligence has to ask something sharper: what does the vendor's system do on its own, in the gaps between the moments a human actually looks at it, and how much of that activity can the bank see in real time, not after the fact? Outsourcing the execution doesn't outsource the accountability. An AI model accessed through a third-party API is still a third-party relationship, and the FFIEC framework still expects the bank to govern it directly, not take the vendor's word for it.

A joint survey from the Bank of England and the FCA found that 46% of firms had only a partial understanding of the AI technologies they were actually using, largely because those technologies came bundled inside third-party models. Meanwhile, 84% of those same firms said they had an accountable person assigned to their AI framework. Put those two numbers side by side and the problem writes itself: accountability lands on a name on an org chart well before that person actually understands what the system is doing. That's a bad combination once the system in question is the one initiating payments.

The new due diligence questions a bank must answer before onboarding a transaction-executing AI vendor

Data security, financial stability, concentration risk: all still on the checklist, all still necessary, none of it sufficient anymore on its own. A vendor executing live transactions needs a sharper, narrower set of questions stacked on top.

What can the agent do without a human checkpoint, and where are the hard stops built in? How does the vendor version its model, especially once post-deployment learning changes the agent's behavior between one audit and the next? Does the audit trail capture every action the agent considered and every path it rejected, along with confidence levels and escalation triggers, or does it only log the transactions that actually went through? What happens when execution fails mid-transfer? Is there an actual rollback mechanism, or just a support ticket and an apology email. SOC 2 is a reasonable baseline here, since it means an independent party actually looked at the controls, but it's a floor, not a finish line, for a vendor touching live banking rails.

Concentration risk cuts sharper in this context too. If the same vendor, or the same underlying model provider, sits behind multiple risk functions at once, a single failure mode can corrupt both the execution and the monitoring of that execution at the same time. Picture a bank where first-line operations, second-line risk, and third-line audit all lean on the same model provider or the same internal AI platform. That's not three independent checks anymore, no matter what the org chart says. That's one blind spot wearing three different badges.

Shadow AI makes the whole picture worse. Unsanctioned tools brought in by employees or third-party partners can skip onboarding entirely, reaching banking systems and transaction records without ever passing through an approved control. And due diligence can't stop at the vendor whose name is on the contract. It has to reach into that vendor's own chain, the sub-processors and upstream model providers the bank never directly touched but whose failure still lands straight on execution.

What transaction-level controls must be built into the vendor relationship from day one

Controls belong at the level of the agent itself, not just floating in the vendor relationship above it. Transaction limits, escalation rules, confidence thresholds, complete action logs: all of it needs to be something the bank can configure, not a fixed setting the vendor decided on somewhere else and shipped as default. A vendor that won't hand over that configuration control is telling you something about how it views the relationship, and it's worth listening.

The three-lines-of-defense model still holds up here, adapted for the job. Line one lives inside the agent: hard caps on transaction size and type, mandatory escalation triggers, action logs that record declined paths as well as completed ones. Line two is the independent risk function, reviewing model performance, digging into escalated cases for patterns, and holding real authority to suspend the agent when something looks off. Line three is audit, and audit needs to reconstruct every decision the agent made straight from the log, not from a summary the vendor decided was worth handing over.

Speed cuts against safety in a way that's easy to underestimate. Speed cuts the window for catching problems: duplicate payments can slip through unnoticed and scams can complete before a warning sign has time to surface, simply because the faster money moves, the shorter the window anyone has to catch a problem. CrowdStrike's 2025 data adds a sobering number: average attacker breakout time stands at 29 minutes, which sits well inside the detection-and-response window most banks are still built around. Controls that only react after the fact aren't fast enough anymore. They need to be pre-configured before the agent ever touches a live transaction, not tuned after the first mistake.

Prompt injection deserves its own line item, not a footnote. F5's research describes a scenario where a threat actor exploits a prompt injection flaw in an AI chatbot to trigger unauthorized transactions, with the missing funds only discovered after several accounts have already been hit. That's not a cybersecurity incident filed away in the IT department. That's a payment fraud vector, and the control architecture needs to treat it as one. Banking APIs that support agentic AI are attractive targets too, especially as open finance expands how many third-party connections touch those APIs. API security governance belongs written into the vendor contract itself, not buried inside a general IT security policy somewhere.

Contracts need to spell out real-time monitoring access for the bank, incident notification windows shorter than a standard vendor SLA, and liability allocation that says plainly what happens when an agent-initiated transaction causes actual harm.

Audit trails and ongoing monitoring when the agent keeps learning

Static model validation checks a system at one point in time. An agent that keeps learning after deployment can drift quietly between one formal review and the next, and a calendar-based check has no built-in way to catch that drift while it's happening.

So what does the audit trail actually need to hold onto? Every action the agent initiated, not just the ones that completed. Every path it considered and rejected, and the reasoning behind rejecting it. Confidence levels and escalation triggers, captured at the exact moment of execution. And the model version active for each individual transaction, which matters a lot once a vendor starts pushing updates on its own schedule, not the bank's.

Deutsche Bank's TPRM AI system, put into place in December 2025, offers a design principle worth borrowing here, even outside the specific TPRM use case it was built for. Every suggested outcome comes with citations back to the exact source passages behind it, a human assessor makes the final call, and the system's accuracy runs around 90% against human validation. Transparency and human decision rights got built in from the start there, not bolted on after something went wrong.

Monitoring cadence needs to account for drift directly: re-validation triggers defined by actual shifts in the agent's decision patterns, not just a date circled on a calendar. NIST's Cyber AI Profile, released in December 2025, pulls together thousands of regulatory expectations from the Federal Reserve, the OCC, the FDIC, and others into one diagnostic framework for AI cybersecurity. Banks can lean on it to structure ongoing monitoring instead of reverse-engineering each regulator's guidance one at a time.

Human oversight still matters to the people running these systems day to day. F5's State of Application Strategy research found that 63% of banking and financial services leaders say human oversight remains central to their security decisions. Audit trail design should reflect that by making human review of agent behavior a routine habit, not something that only happens after a problem surfaces. KPMG's 2025 Banking Survey found that 70% of banking executives are stepping up cybersecurity efforts specifically in response to technologies like generative AI, a reasonable signal that ongoing monitoring needs to include adversarial testing of the agent itself, not just a performance scorecard sitting in a slide deck.

How agentic AI can also improve the TPRM process itself

Manual TPRM is slow, and it's worth being blunt about how slow. A standard document review takes roughly 30 minutes per piece of evidence. A complex, multi-file vendor submission can eat up around three hours of a reviewer's day, according to Deutsche Bank's 2026 reporting on its own process. That's a lot of hours spent reading, and not much of it spent actually thinking about risk.

Deutsche Bank's TPRM AI system, the same one rolled out in December 2025, runs multiple agents in sequence to analyze up to five vendor documents in under two minutes, proposing outcomes with source citations attached before handing everything off to a human for the final call. That's a workable template for human-in-the-loop TPRM at scale. IDC's March 2025 analysis backs up the broader point: agentic AI can strengthen vendor selection, due diligence, contract management, ongoing monitoring, and incident response across the whole TPRM lifecycle, not just at the moment a vendor first gets onboarded.

There's a recursive twist worth sitting with. Once a bank starts using agentic AI to manage its own vendor risk process, that internal system becomes something the bank's own TPRM framework has to govern too. The tool and the target start to blur into each other. What keeps the whole thing honest is the same principle a bank should demand from its transaction-executing vendors in the first place: decision rights stay with people. Agents speed up analysis and surface evidence faster than any human reviewer could alone, but they don't get to replace institutional judgment. A bank that lets them has skipped the actual point of the exercise.

For community banks and credit unions running lean risk teams, this matters even more. AI-assisted monitoring is a real path toward meeting the "proportionate but present" standard regulators are asking for, without needing to triple the size of the risk department to get there.

Building a TPRM framework that can grow as agentic deployment expands

Four pieces have to work together here, not in isolation, and skipping one doesn't make the framework leaner. It makes it broken. Pre-onboarding due diligence covers the vendor's execution scope, its audit trail architecture, SOC 2 and other compliance certifications, and its sub-processor chain. Contractual controls cover bank-configurable transaction limits, escalation rules, liability allocation, notification timelines, and rights over model versioning. Ongoing monitoring covers behavioral drift triggers, an adversarial testing schedule, the NIST Cyber AI Profile as a regulatory shortcut, and independent authority for the second line to actually intervene. Audit infrastructure covers full action logs owned by the bank itself, not summarized secondhand by the vendor, with decision trails that can be rebuilt and records tied to the exact model version active at the time.

Sequencing matters more than most banks give it credit for. Governance foundations come before controlled deployment, and controlled deployment comes before scale. Jump ahead to scale before the audit infrastructure actually exists, and the bank ends up holding liability it has no real way to unwind later, no matter how well the pilot went. This is where most agentic deployments actually go wrong: not in the model choice, but in the sequencing.

Smaller institutions face a scaled-down version of the same framework, not an exemption from it. OCC Bulletin 2025-26 allows for proportionality at community banks, but as Comptroller Gould put it, proportionate doesn't mean absent. All four components still apply, just sized to the institution using them.

Vendor selection turns into a governance decision here, not just a procurement one. A vendor that deploys on the bank's existing rails, hands over configurable controls, and delivers a genuinely complete audit trail is structurally easier to oversee than one that leaves the bank to build its own oversight tooling from scratch around a black box. Pick the second kind of vendor and the bank inherits a project it didn't budget for.

One test cuts through most of the noise: can the bank's own compliance and audit teams, without calling the vendor for help, reconstruct exactly what the agent did, why it did it, and which controls were active at that moment? If the answer is no, the framework isn't finished, no matter how many certifications the vendor happens to hold.

Sources

  1. AI Agent Use Case: Third-Party Risk Management for Financial Services
  2. Agentic AI in Financial Services: Regulatory and Legal Considerations
  3. Putting agentic AI to work in third party risk management
  4. Vulnerabilities for Agentic AI Security in Finance and Banking
  5. Agent-to-Agent Finance: Blockchain Payments and Trust Infrastructure for Autonomous AI Agents
  6. How Agentic AI Will Reshape Payments in: IMF Notes Volume 2026 Issue 004 (2026)
  7. councils.forbes.com
  8. Agentic AI in Banking: 2026 Implementation Guide with Real Bank Case Studies, DORA Compliance, and Generative AI Comparison

More in Examiner Expectations