The Chatbot That Invented a Return Policy. Design for That

Language models make things up with full confidence, and any feature you ship on top of one has to be designed around that behavior from the start. Treat it as the default, not the edge case. The question isn't whether your website's assistant will one day quote a return window that doesn't exist, promise a discount nobody authorized, or cite a policy clause it hallucinated whole.

The question is whether that answer reaches a customer before anyone catches it.

Teams building these features tend to split into two camps. One trusts the model and bolts on disclaimers. The other scopes the model so tightly it can barely answer anything. Neither works alone.

The useful work sits in the tension between them, and it shows up section by section in the choices you make about scope, grounding, and what the output layer is allowed to let through.

The Fabricated Policy Is Already a Legal Problem

The case that keeps getting cited is Air Canada. A grieving passenger asked the airline's website chatbot about bereavement fares. The bot invented a refund policy out of thin air, the passenger acted on it, and when the airline argued in tribunal that the chatbot was a separate entity whose words shouldn't bind the company, the tribunal said no.

The Manatt writeup of the ruling is worth reading in full, because the reasoning is boring in a way that should frighten anyone shipping an LLM feature: the bot was on the airline's site, so the airline owned what it said.

That's one lawsuit. The quieter cost is the small-dollar errors the public never hears about. Wrong shipping windows. Imagined coupon stacking rules.

A promise that a product is in stock when the warehouse says otherwise. Each one trains a customer to distrust you. Enough of them and the feature you launched to deflect support tickets becomes a source of them.

Scope Is the First Guardrail, and the Hardest One to Hold

The first design decision is what the model is allowed to answer at all. The permissive approach lets it try anything and leans on prompts to keep it honest. The restrictive approach limits the model to a small set of intents and hands anything outside that set to a human or a canned response.

The restrictive approach wins almost every time on a transactional site, and it loses almost every time on a general-knowledge one. Pick which site you're running.

Scope is a classifier that runs before the model does, deciding whether the question is in bounds, plus a fallback path for everything else. The practical craft of designing around the ways a client's LLM will fail is mostly this work: naming the handful of questions the feature exists to answer, and refusing the rest cheerfully. A narrow assistant that answers five things well beats a general one that answers fifty things unreliably.

Grounding in Retrieval Versus Trusting the Model's Memory

Once you've decided what the feature answers, the next tension is where the answer comes from. The model's parametric memory is fast, fluent, and wrong often enough to matter. Retrieval against your own documents, product database, or policy pages is slower, uglier to build, and dramatically more defensible.

Research on retrieval augmentation has shown for years that grounding a conversational model in retrieved passages reduces hallucination compared to letting it answer from training alone.

Grounding isn't a cure. A retrieval system can still return the wrong passage, and the model can still ignore what it was given and fabricate anyway. The design question is which failures you'd rather debug.

With retrieval, you can inspect what was fetched, what was passed in, and what came out. Pure generation gives you an answer and a shrug. Pick the architecture whose failures you can see.

Output Guardrails Catch What Scope and Grounding Missed

The third layer is the one most teams skip. Even with scope locked down and retrieval wired up, the model will sometimes produce an answer that shouldn't ship. Output guardrails are the checks that run between generation and display: pattern matches for prices and dates, a second model that verifies the answer against the retrieved context, hard rules that block any response containing a dollar figure the system didn't retrieve.

The permissive instinct is to render the answer and let the user flag problems. The restrictive instinct is to refuse anything the verifier isn't confident about. On high-stakes questions, policies, refunds, legal eligibility, the restrictive instinct wins, and the right UX is a graceful handoff: "I can't confirm that. Here's a link to the policy page, or I can connect you with a person."

On low-stakes questions, product descriptions, store hours, generic how-tos, the permissive instinct wins, because the cost of a wrong answer is small and the cost of constant refusals is a feature nobody uses.

The teams who ship these features well treat the model as a component, not a product. The product is the scope, the retrieval layer, the verifier, and the fallback, with the model doing the one thing it's good at inside that scaffolding.

Build it the other way around and the first fabricated policy will reach a customer. Then a tribunal. Then your legal team.

Latest articles

Related articles