Skip to main content

Set up a human review

Some judgments don't belong to a model. A review sends a dataset to people, collects ratings against a rubric, and shows you where they agree, where they don't, and which items scored worst.

  1. Have a dataset and a rubric

    A review pairs a dataset with a rubric. The dataset holds what gets rated; the rubric is the criteria raters score against. Start a rubric from a preset if you don't have one.

  2. Launch the review

    From the dataset, launch a new review and pick the rubric. The rubric is snapshotted onto the review at this moment — editing the live rubric afterward won't reshape ratings already in flight.

  3. Send the work

    Invite teammates by email, or generate a share link for guest raters — they identify themselves on the rating form and don't need a workspace seat. Each rater scores every item on every criterion.

  4. Read the results

    The review overview gives you more than an average:

    • a coverage matrix (rater × item) — who's rated what, where you're thin;
    • a highlights band — the lowest-scored items and the ones with the most rater disagreement (high spread is a sign the criterion is ambiguous or the case is genuinely hard);
    • a per-item breakdown with mean / min / spread and the rater list.

    Use it to decide what to fix before the next version ships, then mark the review complete.

2 min read