Set up a human review
Some judgments don't belong to a model. A review sends a dataset to people, collects ratings against a rubric, and shows you where they agree, where they don't, and which items scored worst.
- Launch the review
From the dataset, launch a new review and pick the rubric. The rubric is snapshotted onto the review at this moment — editing the live rubric afterward won't reshape ratings already in flight.
- Send the work
Invite teammates by email, or generate a share link for guest raters — they identify themselves on the rating form and don't need a workspace seat. Each rater scores every item on every criterion.
- Read the results
The review overview gives you more than an average:
- a coverage matrix (rater × item) — who's rated what, where you're thin;
- a highlights band — the lowest-scored items and the ones with the most rater disagreement (high spread is a sign the criterion is ambiguous or the case is genuinely hard);
- a per-item breakdown with mean / min / spread and the rater list.
Use it to decide what to fix before the next version ships, then mark the review complete.
2 min read