Check whether the outcome improved
Set agreement checks
Require evaluator pool, independent scores, rubric version, evidence view, task sample, dimension breakdown, disagreement threshold, and fresh-case check for each comparison. Define how missing judgments and adjudication are reported explicitly.
Review slices and cases
Compare agreement by dimension, task difficulty, evaluator pair, and clear or borderline item. Inspect rationales for hidden divergence. Confirm that a shared total does not conceal opposite labels on a consequential dimension.
Approve comparison use
State which agent comparisons the agreement evidence supports and where calibration is required. Record owner, date, sample, and retest trigger. Mark results inconclusive when evaluators saw different context or the sample is too narrow.
What to carry forward
The verification record should show independent labels, evaluator context, dimension slices, disagreement examples, and scope limits. Approve comparison use only within the tested scoring process. Reopen it when evaluator pool, rubric, evidence, or task mix changes before the next comparison with the current dataset.
Technical background: Google DeepMind evaluation research.
Keep the decision with the work.
Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.
Create an account See the product workflow