Skip to main content
DraftNot edited yet. Off the index and not indexed by search.
Guides

How to calibrate an LLM judge against human ratings

Reading time
10 minutes
Assumes
You iterate on prompts and read the output yourself
Updated
Sep 6, 2026

Where this actually starts

You are building a support assistant. It handles billing, refunds, cancellations and account access, and you are three weeks into getting it right.

The loop looks like this. Run some conversations. Paste the ones that went wrong into a doc. Under each, write a line about what it should have said. Hand the doc to Claude, which rewrites the prompt or moves a tool call. Run it again, read the new conversations, and decide whether it feels better than last time.

That loop works, and it forgets. Each round starts from whatever conversations you happened to run that morning. The refund case you fixed on Tuesday is not in this morning's batch, and nothing tells you the rewrite broke it.

The other common version of this is a spreadsheet. One column holds the question, the next holds the answer the assistant gave, and two or three more hold notes from whoever is reviewing. The dev team reruns the questions in a branch, pastes the new answers into a fresh column, and the stakeholders read down the sheet again.

That one is further along than the doc. The cases persist, several people can work in it at once, and the rerun compares against something. What it costs is a person rerunning and pasting every round, which sets how often the loop can turn.

Your notes are already the test set

Look at what is in that doc. Each entry has a conversation that went wrong and a sentence about what a good reply would have contained.

That is a test case. Input, and a standard. The hard part — deciding what good looks like for this specific awkward case — is the part you already did, at the moment you were annoyed enough to write it down.

What is missing is that the doc is a log rather than a set. Entries go in, get used once, and scroll away. Keeping them as a set instead means every round can run against everything you have learned so far, not just against this morning's conversations. The practical version of that is a dataset — the same notes, kept in a shape you can rerun.

The single-question assumption

Both the doc and the spreadsheet push you toward one question and one answer, because that is what fits in a row.

Your assistant does not work that way. A customer gives an order number, asks about a refund, mentions the purchase date four turns later, and pushes back once. The failures that cost you are in that sequence: the number dropped at turn five, the constraint from turn one forgotten at turn six, the reply that was fine on its own and left the conversation nowhere.

A row cannot hold that. So the cases that go into the sheet are the ones that fit the sheet, and the shape of your evaluation starts to be decided by the tool rather than by the product.

A set worth keeping holds whole conversations, scripted exchanges where you write the customer's side in advance, and cases where something plays the customer with a goal and adapts to whatever it is told.

The point at which reading them yourself stops working

Sixty cases, and you ship a change most days.

Rereading sixty conversations after every rewrite takes a couple of hours, so it happens on Fridays, then every other Friday, then when something feels off. Meanwhile Claude is rewriting the prompt daily against the three cases you mentioned this week.

This is the specific failure worth naming. When the thing making changes is faster than the thing checking them, the checking stops being a gate and starts being an occasional audit — and the audit finds problems that shipped a fortnight ago.

What a judge does

A judge is a model handed one output, your rubric, and usually the input that produced it. For each criterion it returns a score and a short reason. Sixty cases take a couple of minutes, so the check can run on every change rather than on Fridays.

What it does not do is decide what good means. It applies a standard you wrote, and it inherits whatever that standard leaves ambiguous. Ask a judge to score "professional tone" and it will produce a confident number for something you have not actually defined.

Can you apply your own standard twice?

Before automating the scoring, find out whether the standard holds still.

If somebody else is also reading output, take twenty cases you have both scored and compare, criterion by criterion. If it is only you, score ten cases, leave them a day, and score them again without looking at the first set.

Landing within half a point is a standard. Two points apart on tone means two standards with one name, and there is nothing for a judge to be calibrated against.

Where the scores diverge, open the case and find the words you read differently. On the support assistant it was "professional tone" — one reading was no contractions, the other was no blame. Rewrite the clause and rescore those ten.

This is worth doing whether or not you ever run a judge. A standard you apply differently on Tuesday and Thursday is producing numbers that move on their own.

What calibration is

Calibration is the check that the judge's scores land where yours land.

It happens per criterion. A judge can match your overall average while running generous on tone and harsh on factual accuracy — the two errors cancel, and the headline number looks correct.

Freeze a sample and score it blind

Fifty to a hundred cases from the set you have. Include the ones you went back and forth on: the day-31 refund, the cancellation that arrived the day after renewal, the reply that was correct and left the customer with nothing to do. A sample of clear-cut cases produces agreement between any judge and any person, because clear-cut cases are where everyone agrees anyway.

Score them by hand before running the judge, and without the judge's numbers visible. Seeing a machine score first changes what you write down, and it moves you toward the machine.

Save the sample. You will run it again after the next model change.

Compare, one row per criterion

CriterionYouJudgeGap
Factual accuracy4.14.0−0.1
Policy adherence3.83.9+0.1
Tone3.44.3+0.9
Says what happens next3.93.7−0.2

Three criteria are close. Tone is nearly a point generous. The overall averages differ by a fifth of a point, so a single blended number would have passed this judge.

Sort the cases by the size of the gap and read the top ten. On the support assistant the tone gap came from replies opening with an apology — marked down by hand as over-apologetic, with nothing in the rubric saying so. Adding a line about apologies closed most of the gap. Most of the gaps you find will point at the rubric.

Where this is manual today

Caliper holds the review scores and the judge scores against the same dataset and rubric. It does not yet compute the agreement between them. Export both and build the table above by hand — an hour, once.

What counts as calibrated

Aim for the judge to sit inside the spread you already produce yourself.

If two passes of your own scoring land within half a point, a judge within half a point of that is doing the job. Holding it to a tighter standard than you meet means fitting to noise.

Where a judge sits outside that band on one criterion, look at the criterion first. A standard a person applies consistently and a model does not is usually one that depends on context the model never received — the support assistant's tone criterion assumed you knew which customers were on a legacy plan.

When it expires

Three things end a calibration: a different judge model, an edit to the judge prompt, and any change to the rubric.

Model swaps are the one worth watching. A newer model can score higher on every published benchmark and read your rubric differently, because your rubric is not a benchmark. Rerun the saved sample and compare the per-criterion gaps against last time.

Reading output yourself does not stop once a judge is running. It moves to a handful of cases on a schedule, which is the same sample and the same comparison as above.

Before you trust the scores

0 of 6 checked