Evaluation
Rubric
The written criteria an output gets scored against, used by both human reviewers and judge models.
Also written rubrics, scoring rubric, criteria.
A rubric is a short list of criteria, each with a name, a description of the behavior it checks, and a scale. Three or four criteria is a normal size. Eight is more than a reviewer will apply carefully, and the last ones end up scored by pattern-matching the earlier ones.
Criteria work when they name behaviors rather than qualities. "Accuracy" is a quality, and two people will fill it with different contents. "Every factual claim is supported by the provided context, and claims outside it are marked as unverified" is a behavior, and two people reading it will usually agree.
The same rubric drives human review and judge runs. That is what makes the two sets of scores comparable, and it is why a vague rubric shows up first as reviewers disagreeing with each other and second as a judge nobody trusts.
Rubrics get snapshotted onto a run in most systems, so editing one later does not reshape results you already have.