SOC 2 Type II Certification Scope for AI Banking Vendors
Scope decisions in SOC 2 audits determine what actually gets tested for AI banking vendors.

SOC 2 Type II is an auditing framework with scope decisions baked into it, and those decisions (which criteria get tested, which systems count as "in scope," how deep the auditor actually looks) decide whether the report tells a bank anything useful. It is not one standard applied the same way to every vendor. For AI banking vendors specifically, that scope question is the whole ballgame.
The standard itself comes from the AICPA, and an independent auditor checks how a vendor protects customer data against a defined set of criteria. The Type II part matters more than most people realize: it means the auditor watched the controls operate over a real stretch of time, not just checked that a policy document existed on the day of the visit. Type I only confirms controls exist at a single point, and that's a weaker claim than it sounds, because a control that exists on paper and a control that actually holds up under daily operational pressure are different things. A report more than twelve months old describes a control environment that may have already changed.
Banks reach for SOC 2 because AI vendors in this space touch account numbers, transaction histories, personally identifiable information, payment credentials (the exact categories that trip regulatory wires and internal risk thresholds at the same time). But here's the structural catch that the rest of this piece is going to unpack: two vendors can each wave a SOC 2 Type II report, and their control environments can look nothing alike. One report might cover exactly what a bank needs to know, while another might leave out the one thing that actually matters. The credential alone tells you almost nothing. Scope tells you everything.
How the Trust Services Criteria are structured and which ones actually matter for AI banking use cases
SOC 2 rests on five Trust Services Criteria: Security, Availability, Processing Integrity, Confidentiality, and Privacy. Security is the only one that's mandatory, the floor everyone has to clear, and the other four are optional add-ons that a vendor has to explicitly choose to include in the audit scope. Nobody gets Confidentiality or Privacy by accident.
For AI vendors operating in banking, Security and Confidentiality are the pair that actually earns the name "essential." Security covers the basics you'd expect: encryption in transit and at rest, role-based access controls, ongoing monitoring for anomalies. Confidentiality governs something narrower and, honestly, more relevant to financial data specifically: how sensitive information gets handled, shared, and eventually retained or destroyed.
Availability starts to matter once a vendor sits inside a time-sensitive process. Payments processing, real-time fraud flagging, anything where a delay isn't just annoying but actually breaks the transaction. A vendor whose system needs to be up and responsive in the moment should have Availability in scope.
Processing Integrity is the one that gets misread the most. It speaks to whether a system processes data completely and accurately, which sounds like it should cover AI decision quality. Processing Integrity checks that data moved through the pipeline correctly; it says nothing about whether the model's output was right, fair, or explainable. That distinction trips people up constantly, and it's worth sitting with for a second: a system can pass Processing Integrity with flying colors while still producing a biased or nonsensical model output, because the criterion was never built to look at that layer.
Privacy comes into play when a vendor is processing personal information under consumer protection rules, and that's increasingly relevant given active CFPB attention to AI-driven consumer interactions.
So what should a bank actually ask a vendor holding a SOC 2 Type II? Which criteria did the auditor test, and do they line up with what this vendor will actually be doing? A vendor certified on Security alone (no Confidentiality, no Processing Integrity) is running a narrower audit than the label implies. The badge looks the same either way. The substance underneath it doesn't.
The four things to actually read in a SOC 2 Type II report before trusting it
Start with the observation period. Twelve months is the clean signal, since it covers a full operating cycle, including whatever seasonal spikes or one-off events put real stress on the controls. Shorter windows aren't disqualifying, especially for newer vendors who haven't had a year to accumulate, but a shorter period should be named and explained, not glossed over.
Next, check which Trust Services Criteria are actually listed in scope. Not implied, not assumed, listed. Match them against what this vendor will actually be doing for the bank.
Then look at the exceptions and how management responded to them. Here's a place where intuition runs backward: a spotless report with zero exceptions isn't automatically more trustworthy than one with a handful of documented exceptions and a solid remediation plan. Auditors find things, and a vendor that names the issue and shows a credible fix demonstrates a level of operational maturity that a suspiciously clean report can't prove either way. An exception in a sensitive control area that never got addressed is what should worry you. That's a finding worth escalating, not a formality to skim past.
The fourth item is the one that matters most for AI vendors specifically: subservice organization carve-outs. A carve-out means the subservice organization (say, an inference provider or a cloud host) sits outside the audit boundary entirely, and the auditor never touched it. If a bank wants assurance on that piece, it has to go get it separately. A carve-in means the opposite: the subservice organization sits inside the tested boundary, and the auditor's work actually covers the full stack.
For AI vendors, the inference layer is where the newest and least understood risk tends to live, and a vendor that carves its model provider out of scope has left a gap that the bank now has to close on its own. Leaving that gap open is no small matter. The leading practice worth looking for: vendors that carve their inference providers into the security boundary, and pair the primary report with each provider's own independent certifications, so nothing sits in a blind spot.
Where SOC 2 stops and the AI-specific risk begins
SOC 2 was built to evaluate data-security controls. It has no criteria at all for model behavior, output integrity, or how autonomously an AI agent acts, and that's simply outside what the framework was designed to measure.
Here's the gap stated plainly: SOC 2 does not validate model accuracy, output quality, or bias. A vendor that points to its Type II report when asked a model-validation question has answered a different question than the one being asked.
What does that mean day to day for a bank? A vendor can hold a spotless SOC 2 Type II and still produce biased credit decisions, unexplainable outputs, or a model whose performance quietly degrades over months. None of that shows up in the audit, because none of it was ever being tested. ECOA and Regulation B don't care whether a human or an algorithm made the credit decision: banks still owe applicants specific, accurate reasons for adverse action, and model opacity isn't a defense regulators accept.
The gap widens further once agentic AI enters the picture. Traditional SOC 2 assumptions lean on a human overseeing the action, a system processing and storing data, and controls preventing unauthorized access. Agentic AI breaks that assumption at the root: the system interprets a goal, plans a multi-step sequence, and executes actions, including financial transactions, without a person in the loop at every step. None of the five Trust Services Criteria say anything about bounding agent autonomy, defining escalation triggers, or limiting what an agent can execute on its own.
SOC 2 Type II, then, gets a vendor through the data-security gate. It says nothing about the model risk gate, the fair lending gate, or the agentic governance gate. Those still lie ahead, unexamined by the audit a bank is holding in its hand.
What regulators currently require beyond SOC 2 (and the gaps they have not yet closed)
The OCC's third-party risk lifecycle guidance, laid out in Bulletin 2023-17, tells banks to scale their diligence to the risk of the activity itself. A SOC 2 review fits cleanly into the due diligence and ongoing monitoring stages of that lifecycle, but it says nothing about contract terms, model-change notification obligations, data residency requirements, or what happens to bank data when the contract ends. Those are separate conversations the audit doesn't touch.
The bigger structural gap sits with model risk guidance. The revised interagency guidance, SR 26-2, issued in April 2026, carried forward the framework from SR 11-7 but explicitly excluded generative and agentic AI from its scope, describing that territory as too novel and too fast-moving to pin down yet. That's a real gap, not a technicality: the most consequential AI deployments happening in banking right now are operating without a definitive federal model risk standard that actually covers them.
Some of that space is getting filled by procurement pressure instead of regulation. In the mortgage and lending world, Fannie Mae's Lender Letter LL-2026-04 requires seller and servicer partners to manage AI and ML risk from their own subcontractors and vendors at a standard no less protective than the letter's own requirements. That's a contractual accountability layer, and it's doing work that a SOC 2 report alone was never built to do.
International frameworks are moving in a similar direction. The FCA's Critical Third Parties regime, in force since January 2025 under PS24/16, requires direct assurance to regulators; a SOC 2 Type II report counts as accepted evidence, but only as one input among several, not the full answer. Under the EU's DORA framework, a current Type II report can reduce the burden of an ICT third-party risk assessment, but it has to be explicitly mapped to the DORA register, and it doesn't stand in for that register on its own.
Every one of these frameworks is sending the same signal from a different angle. SOC 2 is necessary. It's nowhere close to sufficient for what an AI banking vendor actually needs to demonstrate.
How to build the vendor diligence layer that SOC 2 leaves uncovered
This is the layer that SR 26-2 and OCC Bulletin 2026-13 govern, even in the areas where AI-specific rules haven't caught up yet. Banks should be asking AI vendors for model documentation, validation methodology, and evidence of ongoing performance monitoring, not just at onboarding, but on a continuing basis, since model drift is a real operational risk worth revisiting well past the initial launch questions.
The NIST AI Risk Management Framework offers a useful, if voluntary, structure to lean on here. Its four functions (Govern, Map, Measure, Manage) map onto the AI lifecycle and give banks a shared vocabulary for requirements that the Trust Services Criteria simply don't address. Where no regulatory mandate exists yet, NIST AI RMF gives a bank something concrete to frame a vendor questionnaire around instead of starting from a blank page.
Contract language needs to pick up where SOC 2 leaves off. A handful of terms worth insisting on:
- Zero-training-data-use clauses, which prohibit a vendor from using bank or customer data to retrain or improve its own models
- Model-change notification requirements, so the bank gets advance notice before the underlying model changes
- Data residency terms, spelling out exactly where data gets stored and processed
- Termination egress terms, a clear and enforceable process for returning and deleting data once the contract ends
Fair lending exposure deserves its own line of questioning. Banks should require documented fairness testing from any AI vendor touching credit decisions or consumer-facing communication. This isn't theoretical: FCA research published in early 2025 flagged real bias risk in language models used for credit scoring, and that risk doesn't disappear because a vendor's SOC 2 report came back clean.
For vendors whose agents actually execute actions (payments, transfers, account changes) rather than just surfacing a recommendation for a human to approve, a different set of questions applies. What guardrails bound what the agent can do: dollar limits, restrictions on transaction type, defined escalation triggers? What does the audit trail actually capture, every action, every decision point, every escalation? And who inside the bank retains the ability to override, pause, or reconfigure that agent's behavior in real time, not in theory, but as a live operational capability?
What a meaningful SOC 2 scope looks like for a vendor executing real banking transactions
For a vendor that's actually moving money, executing payments, transfers, or account actions, there's a baseline scope worth treating as non-negotiable. Security and Confidentiality both need to be tested, not just Security alone. Availability belongs in scope if the vendor sits anywhere inside a time-sensitive transaction path. And the observation period should run twelve months; anything shorter needs to be named and explained rather than buried in the fine print.
Subservice organizations need the same scrutiny. Inference providers and cloud infrastructure should either be carved into the audit boundary directly, or documented separately with their own independently verified controls sitting alongside the main report. A bank shouldn't have to reconstruct that picture by hand; the documentation should be something a third-party risk team can drop into its own file without extra translation work.
Beyond the report itself, a trustworthy vendor makes more available than a PDF. Configurable controls and explicit guardrails on agent execution should be demonstrable inside the actual product, not just described in a sales deck. Every agent-executed action should leave a complete, immutable audit trail, structured clearly enough that a bank examiner or an internal audit team can reconstruct exactly what happened and why. And the vendor should be able to explain, in plain terms, how model changes get communicated and what options the bank has in response.
What separates a credible vendor from one that's just checking boxes comes down to how deeply accountability is built into the system. The audit trail, the configurable guardrails, the human override capability, these are features built into the system, not compliance paperwork bolted on after the fact. For a bank sitting down to evaluate a vendor in this space, the SOC 2 report is the document you start with, not the one you finish on. What the vendor does after that report, and how it answers the questions the audit never asked, is where the real diligence happens.


