Banks buy chatbots as a customer service investment and govern them as a marketing asset. Both framings miss what the thing actually is.
A chatbot is a channel through which customers make statements to the bank. Some of those statements have legal consequences the moment they are made — and unlike a phone call routed to a trained representative or a letter routed to a dispute department, a chat message lands in a system that was designed to answer questions about branch hours.
That is the central compliance problem with conversational automation, and it gets almost no attention relative to the discussion of what the technology can do.
A customer typing into a chat window may be doing something the institution is obligated to act on. The specific cases worth building for:
An assertion of an error involving an electronic fund transfer. Under the electronic fund transfer rules, a consumer's notice of an error triggers the institution's investigation and resolution obligations, and the notice does not have to arrive in a particular format or through a particular channel. A customer writing "there's a charge on my account I didn't make" in a chat window has, on a reasonable reading, given notice — and the clock is running whether or not anyone at the institution has seen it.
A billing error assertion on a credit account. The credit card billing error provisions contemplate a written notice, and whether an in-chat message satisfies that is a question for counsel rather than for a blog post. The prudent operating assumption is to treat it as one and investigate.
A notice of error or request for information on a mortgage loan. The servicing rules impose acknowledgment and response obligations on qualifying communications, with defined timeframes.
A revocation of authorization for a preauthorized transfer. A consumer may generally revoke by notifying the institution, and an oral or written statement to that effect can be effective. A bot that responds with instructions to contact the originator has not necessarily stopped the clock on the institution's own obligation.
A stop payment request, which has its own timing and documentation requirements.
A complaint. Every prudential expectation for a compliance management system includes complaint capture, categorization, resolution, and trend analysis. A complaint expressed in a chat is a complaint, and a bot that resolves the customer's immediate question without recording it has removed that complaint from the institution's data.
A dispute or expression of dissatisfaction that indicates potential harm, which is the raw material of the trend analysis a compliance function is expected to perform.
The design consequence is unavoidable: the bot must be able to recognize these statements and route them, and the transcript must be captured, retained, and reviewed. Recognition is a functional requirement to write into the specification, not a refinement to add later. Institutions that deploy without it are relying on customers to say things in the channel the institution prefers.
Everything a chatbot says is a statement by the bank. There is no version of this where the bot's answer is the vendor's problem.
Two distinct failure modes, and both are common.
Scripted bots go stale. The bot was configured with the fee schedule, the funds availability policy, and the account terms as they existed at implementation. Then a fee changed, a policy was updated, or a product was retired — and nobody updated the bot, because the change process runs through disclosures, the website, and the branch staff, and the bot was never added to the distribution list. The bot now confidently states a fee the institution no longer charges.
Generative bots invent. A language model asked a question outside its grounding produces a fluent, specific, plausible answer that is wrong. In a banking context that means invented fees, invented eligibility criteria, invented timeframes, and invented policy.
Either way the institution has made an inaccurate statement to a consumer about a material term. Deception under the unfair and deceptive practices standards does not require intent, and a pattern of inaccurate statements is exactly the sort of thing the analysis reaches. Our post on UDAAP covers the standards.
Two controls address this directly.
Content governance. The bot's answer library is customer-facing content about terms and conditions, and it needs the same change control as a disclosure or a marketing piece: an owner, a review cycle, and — most importantly — inclusion in the process that runs whenever a fee, a term, or a policy changes. The practical test: when the institution last changed a fee, was the chatbot on the checklist? For most institutions the honest answer is no.
Hard prohibitions for generative systems. Certain answers should not be produced at all: rate and term quotes, eligibility or qualification statements, credit decisions or predictions of them, tax or legal conclusions, and anything about a specific account before authentication. Refusal with a clean handoff is the correct behavior, and it should be tested rather than assumed.
The rule is simple and the implementations frequently are not: no account-specific information before the customer is authenticated to the standard the institution requires in any other channel.
Related items that get missed:
Session handling on shared devices. A chat transcript that persists in a browser after the customer walks away is a disclosure. Session expiry and transcript handling need explicit decisions.
The transcripts contain nonpublic personal information and are subject to the institution's information security program, its retention schedule, and its privacy obligations. Where the chat platform is a vendor service, that vendor is handling customer information on the institution's behalf — which puts it squarely in the third-party risk program and the information security program's scope, as covered in our privacy and information sharing post.
Generative systems raise a specific question about what happens to the conversation. Whether prompts and transcripts are used for model training, where they are processed, and who can access them are contract terms, and they should be answered before deployment rather than after.
Every chatbot needs a path to a human that is available, obvious, and preserves context. A customer who has to re-explain the situation from the beginning has experienced the automation as an obstacle.
Which leads to the measurement problem in this area. The metric vendors report and institutions repeat is containment or deflection — the share of conversations resolved without a human. It is a bad primary metric, because it counts two very different outcomes identically: the customer whose question was answered, and the customer who gave up.
Metrics that mean something:
Resolution rate, measured by whether the customer's issue was actually addressed — which requires either a follow-up signal or transcript review.
Abandonment mid-conversation, tracked separately from resolution and read as a failure.
Repeat contact rate, where the same customer returns through another channel shortly afterward. This is the most honest single indicator that a "contained" conversation did not work.
Mis-answer rate from transcript sampling. Someone competent reads a sample of conversations and judges whether the bot's statements were accurate. There is no substitute for this and it is the control institutions most often skip.
Legal-trigger detection rate. Of the conversations containing an error assertion, a complaint, or a revocation, how many were correctly recognized and routed. Sampling will find misses, and the misses are the findings.
A frequently unexamined exposure: if the bot touches anything application-related, the credit rules come with it.
Discouragement. Telling a prospective applicant they probably would not qualify, or steering them away from applying, can constitute discouragement under the equal credit opportunity framework — and a bot doing this consistently to a category of inquirers is a pattern rather than an incident.
Adverse action. If any automated interaction results in a denial, the notice obligations attach, including specific principal reasons that must be accurate.
Steering. Directing inquirers toward particular products based on inferred characteristics is a fair lending question regardless of whether a human or a model did the inferring.
The safe design is that the bot provides information and takes applications, and makes no statement about whether a particular person is likely to qualify.
Both, and both governance frameworks apply.
As a vendor arrangement, the chatbot needs the diligence, contract terms, and monitoring in our vendor management post — with particular attention to incident notification, data handling, the right to review, and what happens to transcripts.
As a model, a generative or classification-based system belongs in the model inventory with an owner, a risk tier, and validation appropriate to its consequence, per our model risk management post. The outcomes analysis for a chatbot is transcript sampling, and drift monitoring matters because a vendor's underlying model can be updated without the institution's involvement — which means the system that was tested is not necessarily the system running today. That should be a contractual notification requirement.
Structured coverage is available through Digital Compliance, CFPB Laws: Preventing UDAAP and Other Violations, the Certificate in Deposit Compliance, and the Certificate in Compliance Management System.
A chat interface is a web interface, and the accessibility obligations that apply to the institution's digital channels apply to it: keyboard navigation, screen reader compatibility, sufficient contrast, no reliance on color alone, and no time limits that cannot be extended.
There is also a design point beyond technical compliance. Where chat becomes the primary or only easily reachable support path, customers who cannot use it effectively — including many older customers and customers with limited English proficiency — are excluded from support. Retaining a genuinely reachable phone and branch path is both a service decision and a risk decision.
No recognition of legally significant statements, so error notices and complaints are answered and not routed.
Complaints resolved without being recorded, which removes them from the institution's complaint data and its trend analysis.
The bot omitted from the change process when a fee, term, or policy changes.
Generative answers ungrounded, producing invented fees and eligibility criteria.
Rate quotes or qualification statements permitted, which the bot should refuse.
Account information disclosed before authentication, or transcripts persisting on shared devices.
Containment used as the headline metric, counting abandonment as success.
No transcript sampling, so the mis-answer rate is unknown.
Not in the model inventory, and no notification requirement when the vendor updates the underlying model.
Escalation buried or context-losing, so the human handoff restarts the conversation.
Application-adjacent statements about likely qualification, creating discouragement exposure.
The reframing worth carrying into any chatbot project: this is a regulated communications channel with an automated participant. Everything the institution would require of a trained representative — accuracy, recognition of a dispute, authentication before disclosure, escalation when out of depth, a record of the conversation — is required here too, and the automation does not lower the standard. It just makes the failures consistent and high-volume instead of occasional.
Yes. A customer's assertion of an unauthorized electronic fund transfer, a notice of error on a mortgage loan, a revocation of authorization for a preauthorized transfer, or a stop payment request can each carry obligations that begin when the statement is made, regardless of the channel. The design consequence is that the bot must recognize and route these statements, and transcripts must be captured, retained, and reviewed.
The bank. Everything the bot says is a statement by the institution, and deception under the unfair and deceptive practices standards does not require intent. The two common causes are a scripted bot that went stale after a fee or policy change and a generative bot producing plausible invented answers.
Including the bot in the change process that runs whenever a fee, term, or policy changes. Disclosures, the website, and branch staff get updated; the bot is usually not on the distribution list, so it keeps stating terms the institution no longer offers.
Because it counts the customer whose question was answered and the customer who gave up identically. Better measures are resolution rate, mid-conversation abandonment tracked separately as a failure, repeat contact through another channel shortly afterward, mis-answer rate from transcript sampling, and the detection rate for legally significant statements.
Quote rates or terms, state whether a person would qualify for credit, predict a credit decision, offer tax or legal conclusions, and disclose anything account-specific before authentication. Refusal with a clean handoff that preserves context is the correct behavior, and it should be tested rather than assumed.
A generative or classification-based system does, with an owner, a risk tier, and validation proportionate to its consequence — where transcript sampling serves as the outcomes analysis. Drift monitoring matters because the vendor can update the underlying model without the institution's involvement, which is why notification of model changes should be a contract term.


