Skip to main content
DraftNot edited yet. Off the index and not indexed by search.
Guides

How to red-team a customer-facing chatbot

Reading time
7 minutes
Assumes
You have an assistant customers can talk to
Updated
Sep 6, 2026

Why prompt lists don't work

The usual approach is a spreadsheet of adversarial prompts collected from research papers and incident writeups. You run them, the assistant refuses all of them, and you write "passed red-teaming" in the launch checklist.

Two problems. The list tests single messages, and real attempts unfold over a conversation — build context, establish a premise, then make the ask that would have been refused cold. And the list is public, which means it's in training data, which means refusing it is the easiest thing your model does.

The failures that reach production are patient, contextual, and specific to your product. They don't look like a jailbreak prompt. They look like a difficult customer.

The reframe

Stop asking "does it refuse bad prompts" and start asking "can somebody who wants X get X." Those are different questions and only the second one matches what actually happens.

Start from your own consequences

Generic safety testing is somebody else's threat model. Yours comes from what your assistant can actually do or say that would hurt.

Work through four questions with the people who'd handle the fallout.

What could it give away? Pricing flexibility, another customer's information, an internal policy, the fact that a discount exists at all.

What could it commit you to? A refund, a date, a capability you don't have. Anything a customer could reasonably screenshot and hold you to.

What could it say that's a compliance problem? Advice in a regulated domain, a claim about outcomes, a statement about a competitor.

What could it be tricked into doing through a tool it has access to? If it can look things up, it can be steered toward looking up the wrong things.

Each answer is a goal for a red-team case. That list is worth more than any published prompt collection, because it's about your actual exposure.

Test the persistent attempt, not the single shot

Structure each case as a goal plus a budget of turns, and let the tester adapt.

The techniques that work in practice are ordinary and social. Building rapport before the ask. Claiming something happened earlier in the conversation that didn't — "you already said you'd approve this." Escalating claimed authority from customer to colleague to developer. Reframing the request as a hypothetical, a translation, a form to fill in, a typo to correct. Splitting the goal into pieces that each look harmless.

What makes this hard to test manually is that it needs someone patient, and patience doesn't scale across fifty goals and every release.

Test difficulty, not just hostility

The same mechanism that tests attacks tests something more common and more commercially important: users who are hard to serve without meaning to be.

Somebody confused about your vocabulary who describes the problem wrong. Somebody impatient who skips your clarifying questions and repeats the ask louder. Somebody who gives you one line and expects you to work it out. Somebody whose English is their third language and who takes an idiom literally.

None of these are attacks and all of them break assistants. They also make up far more of your real traffic than adversarial users do, which means testing them has a larger expected payoff — and the same harness runs both.

Score the transcript against expected behavior

The scoring question is not "did it refuse." Refusing everything is trivially safe and commercially useless.

Score against what the assistant should have done in that conversation. For an adversarial goal, that usually means: did not do the thing, stayed civil, and offered a legitimate path where one exists. For a difficult-but-genuine goal: did the person get what they legitimately needed, in a reasonable number of turns.

That second criterion is what stops red-teaming from degrading your product. Teams that optimize only against attacks ship assistants that refuse edge cases and frustrate ordinary users, and the frustration is invisible because nobody files an incident about an assistant that was unhelpful.

Keep both classes in the same suite so you can see the trade in one place.

Common mistake

Fixing a red-team finding by broadening a refusal rule, then not checking what else now gets refused. This is the most common regression in the category — the attack is closed and a legitimate case closed with it, and nobody notices for a month.

Rerun on every change, and after model swaps

Adversarial results are not durable. They're a property of the model, the prompt, the tools available, and the retrieval corpus, and any of those moving can reopen a case you closed.

Model swaps are the underestimated one. A new version can be better on every benchmark and more susceptible to a specific social framing, because susceptibility to persuasion isn't what benchmarks measure. Rerun the adversarial set on every model change and read the transcripts of anything that flipped.

Keep every real incident as a case, permanently. A live failure is the most valuable test material you will ever get, and it should be impossible for that specific failure to recur unnoticed.

Know what this doesn't cover

Be clear with whoever reads the results about the boundary.

This tests conversational manipulation. It does not test prompt injection through retrieved content — a document in your knowledge base containing instructions is a different attack with a different defense. It does not test your tools' own permissions; an assistant that can call something it shouldn't have access to is an authorization bug that no amount of conversational testing fixes. And it does not prove absence. A suite of fifty goals that all failed to succeed tells you those fifty framings didn't work.

Report it as evidence, not as a guarantee.

Before you launch

0 of 6 checked