How to write an eval rubric
- Reading time
- 6 minutes
- Assumes
- You have outputs you need to score
- Updated
- Sep 6, 2026
The failure you won't notice
You write five criteria. You score fifty outputs. You get an average of 3.8 and a chart that goes up over time.
Nobody checks whether a second person scoring the same fifty outputs would have produced 3.8. Usually they would not, and the gap is not small — untested rubrics routinely produce a point and a half of spread between reasonable reviewers on a five-point scale. Every conclusion drawn from that average inherits the spread, including the comparison between last month's version and this one.
This matters more once a judge model is involved, because you will try to calibrate the judge against human scores, and you cannot calibrate against a target that moves depending on who is holding it.
Criteria are behaviors, not qualities
The most common rubric names qualities: accuracy, helpfulness, tone, clarity. These feel like the right list and they are almost unscoreable, because each one is a container into which two readers put different contents.
Replace each quality with the behavior you would point at to justify a score.
Not "Accuracy: is the response accurate?" but "Every factual claim in the response is supported by the provided context. Claims not in the context are marked as unverified." That is checkable. Two people reading it and looking at the same output will usually agree, and where they disagree they can point at the specific clause they read differently — which tells you what to fix.
The test to apply to every criterion: could you show a colleague this text, hand them an output, and predict their score within a point? If not, the criterion is still a quality.
The rewrite that fixes most rubrics
Take each criterion and finish the sentence "I would give this a low score if the response…". Whatever you write is closer to the real criterion than the noun you started with.
Say what each point on the scale means
A criterion with a 1–5 scale and no anchors is five criteria, one per reviewer. Everybody agrees what 5 means. Nobody agrees on the difference between 3 and 4, and most of your outputs land in exactly that gap.
You do not need all five anchored. You need the boundaries that carry decisions. At minimum, describe what earns full marks and what constitutes a failure, and name the specific condition that separates the middle from the top.
Three-point scales are underrated here. If your reviewers cannot reliably distinguish 3 from 4, a scale of fails / acceptable / good produces more usable data than a five-point scale where two of the points are noise.
Keep the list short
Five criteria is a lot. Eight is a rubric nobody completes carefully, because reviewer attention runs out and the last criteria get scored by pattern-matching the earlier ones.
The move that keeps lists short is asking, for each criterion, what decision its score would change. A criterion that never changes an action is documentation, not evaluation. Cut it, or fold it into a criterion that does.
Watch for criteria that are really the same one wearing two names. Clarity and conciseness usually correlate near-perfectly in practice, which means you are paying two reviewer-passes for one signal.
Test the rubric before you trust it
This is the step that separates a rubric from a wish, and it takes an afternoon.
Pick ten outputs, deliberately including the ones your team argued about. Have two people score them independently, without discussing, and without seeing each other's numbers. Then compare.
Read every item where they differ by more than a point. You are not looking for who was right — you are looking for the clause they read differently. Almost every disagreement traces to a specific word doing too much work, and fixing that word fixes the whole rubric.
Rewrite, rescore the same ten, and check that the spread narrowed. Two rounds is usually enough.
Common mistake
Resolving disagreements by discussion and moving on. The reviewers now agree, because they have talked, and the rubric is unchanged — so the next person to use it will disagree exactly as much as the first two did. Fix the text, not the conversation.
Write for the reader who wasn't there
The rubric will be applied by people who were not in the meeting where you wrote it: a new team member, a domain expert from another department, a judge model, or you in four months.
That means no shared context can be assumed. If a criterion references your tone-of-voice guidelines, either quote the relevant part or accept that everyone applies their own version. If it depends on knowing which claims are approved, the approved list belongs in the criterion or beside it.
The same discipline is what makes a rubric usable by a judge model. A criterion that requires context living only in a colleague's head cannot be automated, and the effort of writing that context down is not judge-specific work — it is the work of having a rubric at all.
Before you score anything real
0 of 6 checked