Establish what is happening
Define the agreement unit
Record evaluator population, item sample, rubric version, dimensions, scale, evidence view, and whether agreement means exact score, acceptable range, or shared label. Keep independent judgments separate from adjudicated outcomes for analysis.
Map disagreement shape
Inspect confusion pairs, score spread, dimension reversals, task slices, and rationales. Mark clusters caused by borderline wording, context gaps, evaluator expertise, or real output variation. Preserve the items that reveal the pattern.
Bound the conclusion
State which evaluator group, sample, dimension, and decision the agreement evidence covers. Define when low agreement requires calibration or exclusion. Do not treat high aggregate agreement as proof every dimension is reliable.
What to carry forward
The triage output is an agreement brief with unit, evaluator pool, disagreement slices, causes to test, and conclusion boundary. Stop when a reviewer can see where agreement holds and fails. Keep use provisional where evidence views differ before releasing a new comparison result. from that sample.
Technical background: Google DeepMind evaluation research.
Keep the decision with the work.
Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.
Create an account See the product workflow