Rail Governance

Fair Lending Examination Risk When AI Selects Transaction Parameters

Regulators are catching AI systems that set credit parameters mid-transaction.

Reporter · · 13 min read
Cover illustration for “Fair Lending Examination Risk When AI Selects Transaction Parameters”
Examiner Expectations · September 12, 2026 · 13 min read · 3,037 words

Banks and credit unions have a new kind of credit decision-maker on their hands, and most compliance programs weren't built to watch it. When an AI agent picks a velocity limit, a geographic radius, or a device flag on its own, it has made a credit decision, whether anyone calls it that or not. That decision can produce a discriminatory pattern against a protected class even though the model never saw race, ethnicity, or national origin in its training data. That gap between intent and outcome is exactly what fair lending examiners are trained to find, and it's an operational problem now, one banks and credit unions have to solve before deployment, not after.

Start with what "agentic" actually means in a banking context, because the term gets thrown around loosely. An agentic AI system doesn't just spit out a recommendation for a human to approve. It plans, sets its own intermediate goals, and carries out multi-step transactions on its own, adjusting as it goes. So when a customer starts a transfer, the agent might decide, in real time, what velocity limit applies, how far outside a normal geographic radius counts as suspicious, whether the device being used should trigger a flag, or whether the spending pattern looks "off" relative to history. A credit officer used to set those parameters by hand, in a spreadsheet, with a paper trail. Now they're emergent, chosen by a model mid-session, based on however many features it happens to weigh at that moment.

Sit with the paradox for a second. None of these agents are told a customer's race or national origin. They don't need to be. Proxy variables show up on their own, without anyone engineering them in.

  • Device brand tracks with income level, and income level tracks with race in ways decades of lending research already documents.
  • Geographic radius flags run straight into residential segregation patterns. A flag calibrated on "distance from home address" can quietly turn into a redlining tool.
  • Grocery shopping timing and location correlate with neighborhood demographics, which correlate with the protected classes those neighborhoods are made up of.
  • Transaction velocity norms, learned from historical data, can bake decades of unequal access to banking services right into what the model thinks "normal" looks like.

None of this is speculative. The CFPB's 2022 circular made clear that AI models are subject to disparate impact analysis even when protected characteristics never enter the model as inputs. That's the legal anchor underneath everything else here. Scale makes the problem worse, not better. A rules-based system might check five or six variables. An agentic system can weigh a thousand, which means a thousand potential proxy pathways, and it also means that when a disparity does show up, tracing which variable caused it becomes close to impossible using traditional testing methods. The CFPB's Winter 2025 Supervisory Highlights flagged exactly this: underwriting models running complex AI/ML credit scoring, raising disparate impact concerns serious enough that the agency told institutions to tighten their fair lending testing.

Compare that to how parameter-setting used to work. A rules-based system sets a velocity limit and it stays put until someone changes it, on the record, with a change log. An agentic system might adjust that same limit six times in a single customer session, based on features nobody wrote down anywhere. The institution may genuinely not know which parameters its own agent chose, when it changed them, or why. That's not a hypothetical governance risk. That's a documentation gap that already exists, sitting there, waiting for an examiner to ask about it.

What fair lending examination frameworks actually cover when the decision-maker is an agent

ECOA and the Fair Housing Act don't care what technology made the decision. The obligation to explain a credit decision, and to avoid a discriminatory effect, sits with the institution. Not the vendor, not the model, not the data science team that built it.

Disparate impact theory under ECOA's effects test doesn't require anyone to prove intent. A neutral-sounding policy, like "flag transactions outside a 15-mile radius," that ends up disproportionately burdening a protected class still needs justification. That's been true for decades. What's changed is who, or what, is setting the policy in the first place.

CFPB guidance has raised the bar further: AI-based credit decisions require specific, accurate adverse action reasons at the individual level. Citing a score, or naming the model that produced it, isn't good enough. That's a real problem for parameter-setting agents. If an agent adjusts a velocity limit mid-session, weighing more than a thousand features to get there, producing a compliant adverse action notice means knowing which features drove that particular adjustment, for that particular customer, at that particular moment. A global explanation, meaning a description of what the model generally tends to do, doesn't satisfy ECOA. Regulators are pushing toward counterfactual explanations instead: statements like "if this factor had been different, the outcome would have changed." That's a much harder thing to produce on demand than most institutions currently assume.

The Fair Housing Act stays fully in force for residential mortgage lending no matter what happens to ECOA at the federal level, a point that matters more once the next section gets into the Reg B change.

What does an AI-focused exam actually ask for? A few things, consistently:

  • Full documentation of every AI model touching credit decisions: where the training data came from, how the model is built, what validation testing showed.
  • Proof that disparate impact analysis happened, and that less discriminatory alternatives were actually searched for, not just noted as a theoretical possibility.
  • Adverse action notices that reflect the real basis for a decision, not a generic model output pasted into a form letter.
  • Ongoing monitoring records tracking model performance across demographic groups over time, not a single test run once and filed away.

Regulators and researchers have framed this consistently: high-risk models, meaning ones with real consumer impact, including payment threshold decisions, need evaluation for disparate impact at every stage of development, not just at the moment of launch. For credit unions, the NCUA's September 2025 AI Compliance Plan governs how the agency itself uses AI internally. It doesn't create a new AI-specific exam rule. NCUA still handles AI-related fair lending risk through its existing exam frameworks, worth knowing, because "there's no NCUA AI rule yet" isn't the same thing as "there's no exposure."

Why the April 2026 Reg B amendment does not reduce the compliance exposure

The CFPB's final rule, dated April 22, 2026, and effective July 21, 2026, pulled Reg B's effects-test framework out entirely. The agency's position now is that ECOA doesn't authorize disparate-impact claims at the federal level. Read that as a win, and the read is wrong. Federal ECOA exposure shrinking doesn't mean overall exposure shrinks. It just means the exposure moves.

The Fair Housing Act's disparate-impact theories are still fully available for any AI touching residential mortgage lending. GSE contracts independently require fairness testing and bias controls, regardless of what ECOA says. And several states, New Jersey, California, New York, Colorado, and Illinois among them, run their own expansive anti-discrimination laws that reach straight into algorithmic decision-making, Reg B change or not.

New Jersey is the sharpest example. Its 2025 regulations under the Law Against Discrimination codify disparate-impact standards explicitly, including for AI and automated decision-making tools, across every context the law covers, lending included. Those requirements go further than what federal law currently demands on explainability and governance documentation.

This is structural, not a one-off. State attorneys general have been increasingly active on consumer financial protection, and state enforcement activity around AI decisioning has grown more visible in recent years.

Massachusetts already showed what that looks like in practice: a $2.5 million settlement following an investigation into underwriting practices. The findings centered on human discretion layered on top of models, overrides applied without guardrails, and adverse action notices that failed to actually explain the decision behind them. The settlement came with remediation obligations tied to the underlying underwriting and documentation failures. Separately, a federal law enforcement agency reached a $68 million settlement with a Texas lender over discriminatory marketing and predatory lending aimed at Spanish-speaking borrowers, a reminder that intent-based claims are still very much alive alongside impact-based ones.

Net effect: the Reg B change closed one federal door. It left several others wide open, and it pushed enforcement toward a state-by-state landscape that's more fragmented and harder to predict, even as it gets more coordinated behind the scenes. For any institution with customers spread across multiple states, compliance now has to track where the customer sits, not just where the bank is chartered.

How SR 26-2's agentic AI carve-out creates the governance gap examiners are already probing

The Federal Reserve, the OCC, and the FDIC rewrote their interagency model risk management guidance in SR 26-2 (paired with OCC Bulletin 2026-13 and FDIC FIL-15-2026), replacing SR 11-7 from 2011 and SR 21-8 from 2021. SR 21-8 was a narrow BSA/AML supplement, not a broad rewrite of model risk management itself, so SR 26-2 is really the first ground-up reset since 2011. Fifteen years is a long time in model governance terms, and it shows in how much the new framework has to cover at once.

The new framework is intended to update model risk management guidance that had remained largely unchanged since 2011. The new framework's proportionality provisions are intended to account for the size and complexity of the institution.

Here's the catch, and it's the one most compliance teams are getting backwards. SR 26-2 does not provide a prescribed governance structure specifically built for generative and agentic AI. Institutions are expected to apply their existing risk management principles to govern these systems. Read that as relief, and an institution has already made its first mistake.

Being outside SR 26-2's formal scope doesn't mean an examiner stops asking questions about it. AI oversight questions are increasingly part of the supervisory conversation under existing exam frameworks, given that SR 26-2 itself points back to existing risk management principles for governing these systems. The institution still has to show governance exists, even though no rule spells out exactly what form that governance has to take. SR 26-2 itself says as much: existing principles, materiality, ongoing monitoring, effective challenge, should govern any system the document doesn't formally cover.

That leaves a real documentation vacuum. Agentic systems that adjust transaction parameters continuously don't fit the model of periodic, static validation built for traditional credit scoring. And most institutions aren't even at the starting line here. The Ncontracts 2026 TPRM Survey found only 9% of organizations have identified and documented which of their vendors actually use AI. Before an institution can govern anything, it needs to know what it's governing, and right now, almost nobody does.

The GSEs are moving faster than the prudential regulators on specifics. Freddie Mac's Guide Bulletin 2025-16, effective March 3, 2026, lays out AI/ML governance requirements for approved sellers and servicers, covering performance monitoring, bias and fairness controls, and documentation that can survive an audit. Fannie Mae's Lender Letter LL-2026-04, effective August 6, 2026, extends similar expectations to vendors and subcontractors through a "no less protective" standard, meaning the vendor's governance can't fall short of what's required of the lender itself. NCUA's September 2025 AI Compliance Plan, which governs the agency's own internal AI use, illustrates a structured starting point: inventory AI use cases, identify which carry the highest impact, and apply consistent risk management practices to those.

The vendor accountability gap and why the institution bears the regulatory consequence

Here's the structural problem, stated plainly. When an agent embedded inside a vendor's platform sets a transaction parameter on a customer's behalf, the institution eats the regulatory consequence. Not the vendor.

Most vendor due diligence questionnaires got written before AI became a default feature baked into vendor products, and they often fail to address how the AI gets deployed, what data feeds it, or how its outputs get validated. Contracts rarely spell out AI responsibilities with enough precision to hold a vendor accountable when a disparate impact finding traces back to something the vendor built. That's not a minor drafting oversight. It's the reason institutions end up owning failures they never had visibility into.

Three gaps need closing here, and none of them are optional extras.

  • Evaluate vendor AI governance directly. Ask how bias gets tested for, how outputs get validated, what ongoing monitoring looks like. A vendor that can't answer those questions clearly is itself a finding worth writing down.
  • Rewrite the contracts. AI responsibilities need explicit terms: data standards, security obligations, monitoring requirements, and a clause requiring notification when the vendor makes a material change to an AI-driven feature.
  • Extend ongoing monitoring past onboarding. AI systems drift over time. Vendor monitoring programs need AI-specific performance indicators built in, especially for anything touching consumer-facing decisions.

Shadow AI makes all of this worse. Employees reaching for consumer-grade AI tools to handle transaction-related tasks route around procurement, vetting, and the model inventory entirely. That exposure won't show up in any governance document until an exam drags it into the light.

There's also a continuity angle that gets overlooked. If a vendor's AI-powered transaction system goes down, the operational and customer impact lands immediately. Business continuity planning has to account for AI-driven outages specifically, not fold them into a generic "system down" scenario.

The Filene Institute's framing for credit unions gets at the heart of it: vendor and CUSO partnerships might be the most realistic way for smaller institutions to deploy AI at all. But that partnership doesn't move the fair lending obligation anywhere. The credit union still owns the outcome, full stop, no matter whose model produced it.

What a governance structure capable of catching AI-driven disparate impact actually requires

Everything starts with inventory. An institution can't test, monitor, or document a parameter it doesn't know exists. A complete AI inventory needs to capture what each tool actually does, what data feeds it, what decisions it touches, and who inside the institution owns it.

From there, risk tiering matters. Any parameter-setting agent that influences credit access or payment approval belongs in the top risk tier, and the governance controls wrapped around it need to match that classification, not a lighter one.

Disparate impact testing has to get redesigned for systems that don't sit still. Periodic testing, run once a quarter or once a year, misses parameter drift entirely. If an agent adjusts thresholds continuously, the testing needs to run continuously too, or close to it, checked against demographic outcome data on an ongoing basis. Testing belongs at every stage of the development cycle, not bolted on at the end right before launch, consistent with the framing noted earlier. And when testing does surface a disparity, the institution has to show its work: that it searched for less discriminatory alternatives and evaluated them, not just noted the disparity and moved on.

Explainability has to get built into the pipeline from the start, not retrofitted after an examiner asks for it. Global explanations of model behavior don't satisfy ECOA's adverse action requirements. What's needed is local, individual-level explanation, producible for each specific transaction decision, along with an audit trail capturing exactly which parameters applied, at what values, at the moment the decision got made.

Human oversight design matters just as much as the technical pieces. Override policies need to be explicit and written down. The Massachusetts enforcement action is the clearest cautionary tale here: undocumented human discretion layered on top of a model didn't reduce liability. It created it. Configurable controls that let compliance teams cap parameter boundaries, so the agent can't cross a line without triggering escalation, close that gap.

Someone also has to own this. Pacific AI's 2025 Governance Survey found that 75% of organizations have written AI usage policies, but only 59% have a dedicated governance role assigned to enforce them. That 16-point gap, between having a policy and having a person accountable for it, is exactly where examiners tend to look first.

Across the major frameworks, governance programs for AI in lending consistently center on a core set of concerns: bias and fairness, transparency and explainability, data quality, and security and privacy. Miss any one of those four, and the program reads as incomplete by examiner standards, regardless of how strong the other three look.

The core structural requirement, when it comes down to it, is straightforward even if the engineering behind it isn't: agentic AI running on existing banking rails, with configurable parameter boundaries and a complete audit trail at the transaction level, lets the institution keep control over what the agent is allowed to do, and produces the paperwork to prove it when asked.

How institutions can demonstrate examination readiness before the examiner arrives

Examiners fold AI questions into standard audits now, not separate AI-specific reviews on their own schedule. That means readiness has to be a continuous state, not something assembled in the two weeks before an exam gets announced.

A few categories of documentation come up consistently when the request list arrives:

  • Model documentation covering training data sources, architecture, and validation results, for every AI system touching credit or transaction decisions.
  • Fair lending testing records: disparate impact results, less discriminatory alternative evaluations, and the methodology behind both.
  • Parameter-level audit trails showing which thresholds applied, when, and to which customer segments.
  • Adverse action templates, along with the actual process that turns an agent's output into an individual-level explanation.
  • Vendor AI governance assessments, plus the contract language covering AI-specific responsibilities.

Ongoing monitoring is where readiness really gets tested, though. An institution that can point to a continuous record of demographic outcome monitoring, not a single test run once and shelved, is showing the kind of governance maturity examiners actually look for. A one-time test proves a model was fair on the day it was tested. It says nothing about the months after, and examiners know the difference.

Independent verification carries weight here too. Independent audit documentation and similar compliance-readiness materials signal that controls received external review, not just attested to internally. That distinction, self-attested versus independently verified, tends to shape how an examiner reads the entire governance posture in front of them, before a single specific finding even comes up.

Sources

  1. An Agentic AI Primer for Credit Unions
  2. Emerging Risks in Banking: Q3 2026 Update
  3. riskinfo.ai
  4. Fair Lending Update 2026: Disparate Impact and Underwriting Risk
  5. steptoe.com
  6. monitaur.ai
  7. sullcrom.com
  8. ncua.gov

More in Examiner Expectations