Aglet

Verify Evaluator Agreement Before Comparing Agents

Verification asks whether the scoring process can support the comparison being made. Review exact labels, acceptable ranges, disagreement cases, and fresh items. Report agreement with its evaluator group, sample, and dimension limits rather than as a universal quality claim for this evaluator sample.

Check whether the outcome improved

  1. Set agreement checks

    Require evaluator pool, independent scores, rubric version, evidence view, task sample, dimension breakdown, disagreement threshold, and fresh-case check for each comparison. Define how missing judgments and adjudication are reported explicitly.

  2. Review slices and cases

    Compare agreement by dimension, task difficulty, evaluator pair, and clear or borderline item. Inspect rationales for hidden divergence. Confirm that a shared total does not conceal opposite labels on a consequential dimension.

  3. Approve comparison use

    State which agent comparisons the agreement evidence supports and where calibration is required. Record owner, date, sample, and retest trigger. Mark results inconclusive when evaluators saw different context or the sample is too narrow.

What to carry forward

The verification record should show independent labels, evaluator context, dimension slices, disagreement examples, and scope limits. Approve comparison use only within the tested scoring process. Reopen it when evaluator pool, rubric, evidence, or task mix changes before the next comparison with the current dataset.

Technical background: Google DeepMind evaluation research.

Keep the decision with the work.

Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.

Create an account See the product workflow