Aglet

Turn Judge Calibration Into Better Evaluation Practice

Calibration teaches which words, examples, or evidence boundaries make a judgment reproducible. Preserve the disputed case and the reasoning on both sides. Turn that lesson into a practice that detects score instability before averages conceal it before routine score comparisons.

Keep the lesson for the next incident

  1. Save the disagreement case

    Store rubric version, evidence view, task, output, independent scores, rationales, disputed phrase, resolution, and fresh-case result. Include a disagreement that remained valid after clarification so reviewers do not chase forced consensus.

  2. Improve judge practice

    Add the missing anchor, boundary, example, or evidence field to calibration materials. Schedule independent checks on borderline cases. Explain how the change addresses the specific interpretation split found in the case.

  3. Watch for score divergence

    Set a signal such as widening evaluator spread, dimension-specific reversals, missing rationales, or agreement only on easy examples. Assign an owner to sample judgments and revisit calibration when the signal appears.

What to carry forward

The learning record should connect the interpretation split to the revised anchor or evidence practice and recurrence signal. State which judgment became more reproducible. Keep the lesson tied to the tested rubric and evaluator audience rather than claiming universal agreement across the next scoring cycle and review.

Technical background: Google DeepMind evaluation research.

Keep the decision with the work.

Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.

Create an account See the product workflow