Evaluation
Eval
One scored run of a dataset against a rubric, and the setup that produces it.
Also written evals, evaluation, evaluations, eval run.
The word covers two things people usually mean at once. An eval is the pairing — this dataset, this rubric, this thing being tested. A run is one execution of that pairing, producing a score per item per criterion.
The distinction matters when you compare. Two runs of the same eval are comparable. A run against a dataset somebody has since added items to is not comparable to the run before it, even though the eval has the same name.
Scoring comes from people, from a judge model, or from code assertions — deterministic checks like a regex or a schema validation, for the parts of quality that need no judgement.
What an eval is for, in practice, is answering whether a change made things better or worse. That means the interesting output is rarely the headline number. It is the list of items that scored differently than they did last time.