Evaluation
LLM judge
A model that scores another model's output against a written rubric, instead of a person doing it.
Also written judge, judges, LLM-as-judge, model judge, judge model.
The judge receives an output, the rubric criteria, and usually the input that produced it. For each criterion it returns a score and its reasoning. The point is volume: a person can rate fifty items carefully, and a judge can rate five thousand.
What a judge cannot do is establish the standard. It applies a rubric written by people, and it inherits every ambiguity in that rubric. Two reviewers who disagree about what a 4 means will produce a judge that disagrees with both of them.
Judges also drift in one direction rather than randomly. A judge running half a point generous on one criterion will run generous on every item, on every run, and the scores will look entirely normal. Nothing about the output indicates it. Comparing a judge's scores against human scores on the same items — per criterion, not in aggregate — is the check that finds it.
A new model version can change how a judge reads a rubric, so calibration is worth repeating after any model swap rather than treated as a setup step.