Evaluation
Human review
People scoring AI output against a rubric, usually a sample rather than everything.
Also written human reviews, human rating, manual review, raters.
A reviewer sees one output at a time with the rubric criteria beside it, and scores each criterion. The output is per-item scores, an aggregate, and — more useful than either — the items reviewers disagreed about.
The reviewer who matters is often not on the engineering team. Whoever can say whether a support reply is correct usually works in support. Bringing them in tends to require that they not need an account, a seat, or a walkthrough.
Human review is slow and does not scale, which is the reason judge models exist. It is also the only thing that establishes what good means in the first place, so it does not go away once a judge is running. It moves to a smaller sample on a schedule, where it serves as the check that the automated score still tracks somebody's actual opinion.
Review a sample regularly, and a system whose numbers hold steady while the work quietly gets worse will be caught by a person reading it.