Evaluation
Golden dataset
A set of test cases picked by hand, each with a written note about what a good answer would contain.
Also written golden set, golden datasets, eval set, evaluation set, test set.
Usually fifty to a hundred and fifty cases. Each one holds an input and a standard — either the answer you would want, or a description of what any acceptable answer has to do. You run the whole set whenever you change a prompt or swap a model, and compare against the last run.
The word golden refers to the standard rather than the quality of the examples. The cases themselves are often the ugly ones: the complaint, the request that sat exactly on a policy boundary, the question your system should have refused.
Size matters less than composition. A set exported at random from production will mostly contain requests that already work, and those cases pass on every version, so the score barely moves when something breaks. A set assembled from incidents, boundary cases and refusals will move.
A golden dataset is only comparable to itself. Adding or editing items changes what the score means, which is why additions usually happen in dated batches rather than continuously.