How to test a chatbot across a whole conversation
- Reading time
- 6 minutes
- Assumes
- You have an assistant that holds conversations
- Updated
- Sep 6, 2026
What single-turn tests can't see
A question-and-answer test set measures one thing: given this input, is the output good. That's a real property and it's a fraction of what a conversational system does.
Four failure classes live entirely in sequence. Context loss — the customer gave their account number at turn one and the assistant asks for it again at turn five. Contradiction — it says one thing early and something incompatible later. Drift — the register or the constraints loosen over a long conversation, and by turn eight it's agreeing to things it refused at turn two. Failure to converge — every individual reply is fine and the customer still doesn't have what they came for after ten turns.
None of these are visible in a per-message score. All of them are common.
The unit of evaluation
For a conversational system the unit is the transcript, not the message. A transcript where every reply scores well and the customer left without an answer is a failing case.
Two shapes, for two different questions
Scripted sequences. You write the user's turns in advance, in order, and check what comes back at each step. Deterministic, repeatable, and ideal for specific mechanics — does it retain the account number, does it honor the constraint stated at turn one, does it handle a correction at turn four.
Use these when you know exactly which sequence you're worried about. They're precise and they only test the paths you thought of.
Simulated people. Something plays the user with a goal, adapts to each reply, and stops when the goal is achieved or abandoned. Non-deterministic, and it explores paths you didn't anticipate — which is the point, and the reason it finds things scripts don't.
Use both. Scripts pin known mechanics; simulation finds what you missed.
Test difficult users, not only hostile ones
Simulation gets discussed as an adversarial technique, and the adversarial case is the smaller half of its value.
Most of your traffic is people who are hard to serve without meaning to be. Someone confused about your vocabulary describing the problem with the wrong words. Someone impatient who skips your clarifying questions. Someone giving you one line and expecting you to work it out. Someone whose English is their third language taking an idiom literally.
These break assistants routinely and they never appear in a test set written by the team who built it, because the team knows the right words.
Score the transcript, not the last message
Your rubric needs criteria that only make sense over a whole conversation.
Goal achievement. Did the person get what they legitimately came for? This is the criterion that matters most and the one single-turn evaluation cannot express.
Efficiency. How many turns did it take? An assistant that gets there in nine turns when four would do is failing, and no per-message score will say so.
Consistency. Did it contradict itself? Did constraints from early turns survive to the end?
Retention. Did the user have to repeat information they already gave?
Keep your per-message criteria as well — tone and accuracy still matter. The transcript-level criteria are additions, not replacements.
Set a turn budget and treat exhaustion as a result
Every multi-turn case needs a cap, for cost and for termination.
The budget is also a measurement. A case that regularly exhausts its budget without resolving is telling you the assistant can't converge on that goal, which is a finding rather than an inconclusive test. Record exhaustion as its own outcome instead of folding it into failure — "didn't finish in eight turns" and "gave the wrong answer" are different problems.
Set the budget from what's reasonable for a person, not from what the assistant needs. If a competent agent would resolve it in four turns, eight is a generous cap and twenty is hiding a problem.
Accept the non-determinism, then manage it
Simulated runs differ between executions, which makes comparison harder than with a fixed set.
Three habits make it workable. Run the important cases more than once and look at how often the goal is achieved rather than at a single result. Read transcripts rather than only scores, because the interesting information is in how the conversation went wrong. And when a simulated run finds a real failure, convert it into a scripted case so that specific sequence is pinned deterministically from then on.
That last one is the loop that makes this pay off: simulation explores, scripts pin, and the pinned set grows with every genuine discovery.
Common mistake
Judging a multi-turn run by reading the final message. The failure is usually upstream — a constraint dropped at turn three, a question that should have been asked and wasn't. Read the whole transcript, or you'll fix the symptom at the end.
Before you rely on the results
0 of 6 checked