Evaluation
Inter-rater agreement
How closely two people scoring the same outputs against the same rubric land on the same numbers.
Also written inter-rater reliability, rater agreement, rater disagreement.
Two reviewers, the same items, no conversation between them. Compare the scores.
Wide disagreement is a finding about the rubric rather than about the reviewers. A criterion that two careful people read differently will be read differently by everybody else too, including a judge model, and every number built on it inherits the spread.
Untested rubrics routinely produce a point and a half of spread on a five-point scale. That is larger than most of the effects people try to detect with them.
The useful move on a disagreement is not to talk it out. Two reviewers who discuss an item will agree afterwards, and the rubric is unchanged, so the next person to use it disagrees exactly as much. The fix is to find the clause they read differently and rewrite it, then rescore.
Agreement between reviewers also sets a floor for what to expect from a judge. If your own team disagrees by half a point, holding a model to tighter than that is chasing noise.