Aglet

Verify an Agent Judge Calibration

Verification tests shared interpretation before comparing agent scores. Review independent judgments and reasons, then inspect whether anchors predict decisions on fresh cases. Agreement is useful evidence only within the evaluator group, examples, rubric, and context that were calibrated for this comparison.

Check whether the outcome improved

  1. Set calibration checks

    Require rubric version, evidence view, dimension definitions, anchors, ambiguity rule, independent scoring, and rationale. Define acceptable disagreement and a re-review trigger. Keep evaluator identity and task context available for analysis.

  2. Review fresh examples

    Apply the calibrated rubric to clear, borderline, failure, and unseen cases. Compare scores and rationales by dimension. Inspect whether a shared score masks different reasoning or whether a difference is supported by distinct evidence.

  3. Approve the scope

    State which evaluators, examples, dimensions, and decisions the calibration supports. Record owner, date, rubric version, and retest trigger. Mark it inconclusive when the group is small or evidence views are not comparable.

What to carry forward

The verification record should show rubric, evidence views, independent scores, rationales, anchors, and scope limits. Approve calibrated use only for the tested evaluator group and task contract. Reopen it when the rubric, evaluator prompt, examples, or evidence view changes before publishing new scores.

Technical background: Google DeepMind evaluation research.

Keep the decision with the work.

Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.

Create an account See the product workflow