Skip to main content
← Glossary

Evaluation

Expected output

The standard on a test case — either the answer you want, or a description of what any acceptable answer must do.

Also written expected outputs, expected answer, reference answer, ground truth.

Two forms, and the choice depends on how many good answers exist.

For a question with one right answer — a policy, a date, a calculation — write the answer out. A judge compares meaning rather than characters, so the wording can differ, but the substance has to match.

For anything conversational, describe what an acceptable answer has to do. "States the 30-day window, says this request falls outside it, offers the exceptions process." Several replies satisfy that, and several replies are fine.

Writing a single model answer for an open question is the most common cause of frustration on a first eval. Good responses score badly for phrasing things differently, the numbers stop tracking quality, and the judge takes the blame for a problem in the standard.

Expected output is distinct from the answer a system already gave. That one gets rated as-is — see captured output.