Build a useful investigation brief
Collect paired judgments
Save exact rubric version, evaluator prompt or instructions, evidence view, task, output, score, rationale, and timestamp for each judgment. Pair evaluators on the same examples. Mark missing context or post-hoc explanation.
Locate the interpretation split
Compare dimension definitions, anchors, evidence thresholds, and treatment of ambiguity. Classify the split as wording, example coverage, evidence access, scale use, or an actual output difference. Avoid collapsing distinct dimensions into one disagreement.
Test the correction
Revise one anchor, instruction, or evidence view and rerun the disputed cases. Add a fresh borderline case. State whether disagreement narrows, persists for a valid reason, or reveals that the dimension should be separated.
What to carry forward
The investigation brief should link paired judgments, evidence views, interpretation split, revision, and rerun. End with a localized calibration gap or a documented valid disagreement. Do not call agreement improved if only the scoring prompt changed without new examples for comparison.
Technical background: Google DeepMind evaluation research.
Keep the decision with the work.
Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.
Create an account See the product workflow