Rail Governance

PCI DSS Scope Determination for AI Agents Handling Payment Card Data

AI agents handling payment data fall squarely under PCI DSS scope based on capability, not labels.

Correspondent · · 13 min read
Cover illustration for “PCI DSS Scope Determination for AI Agents Handling Payment Card Data”
Accountable AI · September 30, 2026 · 13 min read · 2,952 words

PCI DSS v4.0.1 doesn't say the words "artificial intelligence" anywhere in its text, and it doesn't need to. The standard reached full mandatory status on a tight timeline: v3.2.1 retired March 31, 2024, v4.0 followed it into retirement on December 31, 2024, and by March 31, 2025 the last set of future-dated requirements became binding for every organization that stores, processes, or transmits cardholder data Shuttle. That timeline matters because it means the version of PCI DSS banks and fintechs are building against right now is the full v4.0.1 requirement set, not a transitional draft.

Scope, under this standard, is a function of what a system can do, not what label sits on the box. A system that stores, processes, or transmits cardholder data is in scope. So is a system that could affect the security of the cardholder data environment, even if it never touches a card number directly. That second clause is where most AI deployments quietly step into trouble.

Consider an LLM-powered internal tool that can read transaction logs or trigger a payment reversal. Its system prompt might never mention a card number. Doesn't matter. Capability determines scope, not content, and an agent with that reach is a system component under PCI DSS whether anyone architected it to be one or not.

The indirect case cuts even further. Scope gets assessed against the environment the tool sits inside, not the technology label stapled to it. Practically, that means asking what the tool can reach, where the information it touches actually goes, and what other systems it talks to.

So the frame for anyone building agentic payment tools is blunt: no architecture choice opts an AI agent out of PCI compliance by naming it something else. The only question that ever matters is whether its capability touches the cardholder data environment. An AI model that does not directly hold card numbers is nonetheless in scope if it has access to connected systems that handle cardholder data or can change those systems.

The PCI SSC's statements on AI, as of September 2025

The Council went quiet on AI for a long time, then moved twice in six months https://blog.pcisecuritystandards.org/new-guidance-integrating-artificial-intelligence-into-pci-assessments. First came narrower guidance on March 17, 2025, aimed at assessors: AI can help with document review and drafting work papers, but it cannot assume the role of an assessor Shuttle. That's a small but telling signal, because it establishes a principle the Council would lean on again months later, on the other side of the compliance relationship https://blog.pcisecuritystandards.org/new-guidance-integrating-artificial-intelligence-into-pci-assessments.

New guidance from the PCI Security Standards Council addresses the use of AI in PCI assessments, emphasizing that AI is a tool, not an assessor, with human assessors remaining responsible for all findings and final decisions https://blog.pcisecuritystandards.org/new-guidance-integrating-artificial-intelligence-into-pci-assessments. The guidance acknowledges that payment systems are moving from human-managed operations toward agentic AI with real agency to act on its own behalf, and it says outright that the pace of change makes risk assessment hard. That's a rare admission from a standards body: the target is moving faster than the guidance can chase it.

On substance, the September principles confirm what capability-based scoping already implied Shuttle. Requirement 3, protecting stored cardholder data, and Requirement 4, securing data in transit, apply to AI-based systems exactly as they apply to anything else. No exemption, no asterisk.

The accountability language is where the guidance gets specific. AI has to be deployed so its actions can be logged and monitored, and a named human individual has to be held responsible for those actions. Read that again: not "a team," not "the vendor," a human individual. It's stated as a compliance expectation, which means an assessor can reasonably ask for it by name.

The SSC acknowledged that agentic AI operating as a payment actor requires access to payment card data, but noted that single-use PANs and payment tokens are a path to keeping that data protected. That's a meaningful signal about where the Council wants the industry to land architecturally, and it lines up with what the card networks were already building, covered further down https://blog.pcisecuritystandards.org/new-guidance-integrating-artificial-intelligence-into-pci-assessments.

PCI SSC published formal AI Principles guidance on September 11, 2025, authored by Andrew Jamieson (the first official PCI SSC guidance specifically addressing AI in payment environments), alongside separate March 17, 2025 guidance stating that AI can assist with document review and work papers but "AI cannot assume the role of an assessor" Shuttle. Compliance teams should treat it that way: as the clearest read available on how existing requirements apply to AI, not as fresh obligations waiting to be checked off. As of September 2025, the SSC has not issued AI-specific requirements inside PCI DSS itself, since the document released that month is guidance, not a requirement update, and compliance teams should treat it as interpretive authority rather than as a new checklist.

The four requirements that create the most friction when an AI agent touches the CDE

PCI DSS has requirements, but four carry almost all the weight when an AI agent enters the picture. The friction shows up in different places for each one.

Requirement 3 governs protection of stored cardholder data through encryption, tokenization, or masking, and it applies to any AI system that retains card data anywhere, in its context window, in logs, or worse, in a fine-tuning dataset. That last case is the sneaky one. A model fine-tuned on real support transcripts that happened to contain card numbers has baked cardholder data into its parameters. Good luck redacting that after the fact.

Requirement 6 covers secure systems and software development, and it extends to the AI model's own lifecycle, training pipelines, deployment processes, and model parameters all count as system components. Requirement 7 restricts access by business need, which means it governs who can reach the training datasets that might contain payment records, and who can reach the model parameters derived from them.

Requirement 8 is arguably the sharpest one for agentic systems. Identification and authentication rules, specifically PCI DSS 8.2.2, require every non-human entity to carry a unique identifier, so every action traces back to a specific account. Without that, there's no way to reconstruct an audit trail after something goes wrong. An AI agent acting as a transactor without its own identifier is, from an audit standpoint, a ghost.

Requirement 10, logging and monitoring, is where things get genuinely hard. The requirement calls for comprehensive logs of access to system components, and for AI that means logging the full reasoning chain, prompt inputs, retrieved context, every tool call, every agent-to-agent handoff, and the final output. Not just the answer. The whole path that got there.

Why does that matter so much more for AI than for a traditional system? Because AI is non-deterministic. The same input can produce a different output tomorrow. Models drift. A static log snapshot, the kind that worked fine for a rules-based payment engine, may not adequately represent what the model actually did at the moment it did it, since the same input can produce different outputs and models can drift over time. So the practical move for any compliance team is to map these four requirements against the agent's real architecture before the QSA shows up, not after. Retrofitting Requirement 10 logging onto an agent that's already in production is a much harder job than designing it in from the start.

Leakage points where card data enters AI pipelines that teams overlook

Most of the risk here is mundane, and it happens because nobody built the capture flow to stop it.

Start with training data. If a voice agent hears a customer read a card number aloud, that audio gets processed somewhere. If a chat agent receives a card number typed into a message, that text runs through an LLM. Unless the architecture explicitly blocks it, that data flows straight into systems that were never PCI-certified, and potentially into the training data itself. It isn't hypothetical. Customers type card numbers into web chat, or read them out loud, because the agent asked "what card would you like to use?"; because the question lacks a proper capture flow, it produces this outcome. That's the whole failure mode: a poorly worded prompt with no redirect to a secure channel.

Conversation summaries carry the same risk in a quieter form. Most conversational AI platforms auto-generate session summaries for handoff, QA, or analytics, and a summary line like "used the card ending in 4471" landing in a BI dashboard puts that dashboard in PCI scope, even with no full PAN anywhere in it. Voice transcription doubles the exposure: a speech-to-text pipeline processes raw audio of someone reading sixteen digits, a CVV, and an expiry date before any tokenization has happened, and if that audio or its transcript gets logged for QA or training, the card data is in scope twice over, once in the audio store and once in the transcript.

Multi-agent handoffs introduce a different kind of problem. A booking agent calls a rebooking agent, which calls a refund agent. Session context passes between all three, and those systems may not share the same security controls. If a token, or worse a raw PAN, rides along in that context because it was convenient to pass it that way, every single hop in the chain inherits PCI scope.

Shadow AI deserves its own mention. Developers who deploy AI coding assistants against payment code repositories without proper risk management create Requirement 12.3 violations, and those exposures compound quietly, adding up over time and raising the odds of a breach. Fraud detection systems, meanwhile, are the one AI use case nobody argues about: transaction pipelines carry card numbers (sometimes truncated or tokenized, sometimes not), transaction amounts, merchant identifiers, and cardholder verification data, so they're unambiguously inside the CDE.

Agentic AI makes the whole boundary problem worse, not better. These systems call APIs on their own, orchestrate multi-step workflows, pull from knowledge bases, and hand off tasks to other agents, often without a human approving each step. The CDE boundary was designed around discrete, human-operated components. An agent that autonomously spans five systems in one workflow doesn't respect that boundary. It just runs through it.

Three architectural patterns that keep an AI agent out of PCI scope

The underlying principle is simple to state, harder to build: the AI agent, the LLM, the conversation engine, the decision logic, has to be architecturally separated from the payment infrastructure, with no card data ever crossing that line. The agent decides when to initiate a payment. The payment layer handles the card. That's the whole division of labor.

For voice agents, DTMF capture is the established pattern. The customer punches card digits into the phone keypad, the tones get captured in a PCI-compliant call channel, and they never reach the AI's speech-to-text engine at all. Done right, the agent never touches card data, so its own PCI scope shrinks to almost nothing, and the compliance burden sits with the secure capture and payment layer instead. If the recording-pause gap around DTMF capture resumes even a beat late, it can catch a partial PAN on tape. That's a timing bug, and it needs to be tested in the actual implementation, not assumed to work because the design document said so. The pattern also has a natural limitation: it assumes a live customer is on the line at the exact moment payment is due. It doesn't hold up for an agent trying to rebook a flight after a disruption when the customer isn't there to pick up the phone.

For chat agents, payment links do the same job. The agent generates a link through the payment layer's API, passing amount, merchant ID, and a reference, never card data, and the customer completes payment on a PCI-certified hosted checkout page. The chat platform, the LLM, and every conversation log stay clean. The agent only ever sees a confirmation: paid or not, amount, reference. The catch is that this pattern assumes a fixed price and a single merchant at a moment the customer is actively present, which breaks down for multi-vendor journeys where the price is moving in real time or the agent has to act on its own.

For fully automated agents, tokenized card-on-file closes the gap. The customer stores a card once through a PCI-certified flow, and from then on the agent sends payment requests using a token, never the card itself, while the payment layer redeems that token with the processor. The gateway that stores the card, and the payment layer that redeems the token, stay inside PCI scope. The agent itself sits outside it.

Even a well-built system will occasionally have a customer speak card details anyway, out of habit or confusion. When that happens, the system needs to do three things at once: refuse to treat the spoken data as payment input, strip the card number out of any transcript, and still route the actual payment through DTMF capture. None of these patterns matter in isolation. DTMF masking, tokenization, and link-based capture all need to run consistently across every channel, voice and digital alike, because a single gap in one channel reopens CDE scope for the whole system.

Agent-Scoped Tokenization in Card Networks' Existing Rails

Through 2025 and into 2026, Visa, Mastercard, and Google each shipped a version of the same basic idea: instead of an AI agent holding a raw card number, it holds a token scoped to that specific agent, that specific merchant relationship, and a defined spending policy.

Mastercard moved first with Agent Pay, announced April 29, 2025. It's a framework that lets verified AI agents transact using Agentic Tokens, built as an extension of Mastercard's existing Digital Enablement Service. Launch partners included Microsoft, IBM, Braintree, and Checkout.com. Mastercard followed up on January 27, 2026 with Agent Suite, slated for availability in the second quarter of 2026, aimed at product discovery for banks and conversational shopping for merchants. Alongside it sits Mastercard's Verifiable Intent offering, which implements Know Your Agent frameworks, cryptographic verification meant to tell a legitimate shopping agent apart from a malicious bot. None of this is a bolt-on product line. It's a capability layered onto existing tokenization rails, so most merchants will get it through their processor and most agents through an LLM-platform partnership rather than a direct integration.

Visa's answer, Intelligent Commerce, launched the same month, April 2025. It offers APIs and SDKs purpose-built for tokenization, authentication, and transaction controls in agentic contexts, and its partner list reads like a cross-section of the entire AI industry: Anthropic, IBM, Microsoft, Mistral AI, OpenAI, Perplexity, Samsung, and Stripe.

Google took a slightly different structural approach with its AP2 mandate model, which layers three cryptographically signed mandates on top of a transaction: an Intent Mandate that records what was actually requested, a Cart Mandate that locks in the exact items and amount before payment fires, and a Payment Mandate that ties the verified cart to the payment network. The effect is that what the customer sees is what gets charged, and the agent can only act inside the boundaries of the mandate it was granted.

Agent-scoped tokens structurally align with the SSC's September 2025 preferred direction of single-use PANs and payment tokens for agentic AI, giving deployers using these frameworks a compliance posture that keeps raw PANs outside agent context by design. A deployer building on any of these frameworks inherits a compliance posture where raw PANs stay outside the agent's context by design, not by policy memo. That's a meaningfully different starting position than trying to bolt tokenization onto an agent after the fact.

Requirements for a PCI-Compliant Audit Trail for an AI Agent

PCI Requirement 10 compliance for AI is more demanding than for static systems: a compliant audit trail must capture prompt inputs, retrieved context, tool calls made, agent-to-agent handoffs, and outputs returned (not just final transaction records).

Why does the bar sit so much higher here? Because of that non-determinism problem again: identical inputs can produce different outputs, and the model itself can drift over time, so a log that only records what went in and what came out, without the reasoning in between, can't reconstruct what the agent actually decided or why it decided it.

Logging alone doesn't satisfy the Council's accountability requirement either https://blog.pcisecuritystandards.org/new-guidance-integrating-artificial-intelligence-into-pci-assessments. The September principles are specific: a named human individual has to be identifiable as responsible for each AI-driven action, which means agent identities set up under Requirement 8.2.2 need to map back to an actual person in the organization's accountability chain.

What regulators are straining to observe here is the relationship between the instruction and the action that follows it, not the final transaction, since most payment regimes already handle that fine — an agent-initiated payment may not correspond to any single explicit transaction-level instruction a human gave in the moment.

Banks aren't only answering to PCI on this, either. Financial regulators are moving in parallel. The revised interagency model risk management guidance issued April 17, 2026, SR 26-2 and OCC Bulletin 2026-13, states that existing risk management principles apply to institutions even in areas where generative and agentic AI technically fall outside the guidance's formal scope. It's the same discipline banking regulators are already asking for, arriving from two directions at once. Mandate-based authorization, the pattern seen in Google AP2 and the card network token frameworks, offers a partial answer, in which the human grants a cryptographically signed, narrow authorization and the agent can only act inside it, making the instruction-to-action relationship auditable in principle. Before deployment, compliance teams should build per-agent unique identifiers (Req 8.2.2), structured logs of the full reasoning chain, human accountability mapping, and a process for detecting model drift that could cause the same configured agent to behave differently over time.

Sources

  1. AI Payment Security: How AI Agents Handle Card Data Without Breaking PCI | Shuttle
  2. New Guidance: Integrating Artificial Intelligence into PCI Assessments
  3. AI Principles: Securing the Use of AI in Payment Environments
  4. The PCI Problem with AI Travel Agents and Conversational Booking
  5. PCI DSS for Customer Support AI Agents | PADISO Blog
  6. Agentic AI and PCI DSS Compliance
  7. AI and PCI Compliance: What Every Company Needs to Know in 2026 | Very Good Security
  8. AI in PCI DSS Compliance Assessments: PCI Security Standards Council Weighs In
Filed underAccountable AI

More in Accountable AI