Skip to main content
← Glossary

Evaluation

Golden dataset

A set of test cases picked by hand, each with a written note about what a good answer would contain.

Also written golden set, golden datasets, eval set, evaluation set, test set.

Usually fifty to a hundred and fifty cases. Each one holds an input and a standard — either the answer you would want, or a description of what any acceptable answer has to do. You run the whole set whenever you change a prompt or swap a model, and compare against the last run.

The word golden refers to the standard rather than the quality of the examples. The cases themselves are often the ugly ones: the complaint, the request that sat exactly on a policy boundary, the question your system should have refused.

Size matters less than composition. A set exported at random from production will mostly contain requests that already work, and those cases pass on every version, so the score barely moves when something breaks. A set assembled from incidents, boundary cases and refusals will move.

A golden dataset is only comparable to itself. Adding or editing items changes what the score means, which is why additions usually happen in dated batches rather than continuously.