Glossary
Plain definitions for the words that come up when you build, test and measure AI features.
Evaluation
Captured output
An answer a system already produced, stored on a test case so a person can rate it as it stands.
Eval
One scored run of a dataset against a rubric, and the setup that produces it.
Expected output
The standard on a test case — either the answer you want, or a description of what any acceptable answer must do.
Golden dataset
A set of test cases picked by hand, each with a written note about what a good answer would contain.
Human review
People scoring AI output against a rubric, usually a sample rather than everything.
Inter-rater agreement
How closely two people scoring the same outputs against the same rubric land on the same numbers.
LLM judge
A model that scores another model's output against a written rubric, instead of a person doing it.
Regression
A change that fixed one thing and broke another, usually somewhere nobody thought to check.
Rubric
The written criteria an output gets scored against, used by both human reviewers and judge models.
Measurement
Goodhart's law
When a measure becomes a target, it stops being a good measure.
Guardrail metric
A number you expect to hold steady while you move a different one, named before the change ships.
Pre-registration
Writing down what you expect a change to do, with a number and a date, before any evidence arrives.