Build a useful investigation brief
Assemble independent labels
Save evaluator, rubric version, evidence view, task, output, dimension scores, labels, rationales, and timestamps. Keep adjudicated labels separate. Sample clear, borderline, and disagreement-heavy cases for comparison on the same evaluation items.
Classify the split
Compare dimension definitions, anchors, source context, evaluator instructions, and output details. Mark disagreement from missing evidence, ambiguous rubric, evaluator drift, item difficulty, or a real behavior distinction. Record multiple causes when they interact.
Test a remedy
Apply one change such as a new anchor, shared context, evaluator pairing, or dimension split. Rescore the disputed sample and a fresh sample. State whether agreement improves, remains legitimately divided, or exposes a different measurement problem.
What to carry forward
The investigation brief should preserve independent labels, context, rationale comparison, disagreement class, remedy, and fresh rescore. End with an attributed agreement gap or bounded uncertainty. Do not use adjudication alone as evidence that independent scoring is reliable for this comparison.
Technical background: Google DeepMind evaluation research.
Keep the decision with the work.
Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.
Create an account See the product workflow