Decide where the work belongs
Rank decision risk
Describe what a false positive or negative agreement conclusion could change. Give weight to release gates, broad score rollups, critical task slices, and judgments that are expensive to revisit after publication.
Compare gap leverage
Review whether calibration examples, clearer evidence, evaluator pairing, or dimension separation would resolve the gap. Prefer an intervention that addresses repeated disagreement rather than one exceptional item in the queue.
Choose an agreement queue
Select dimension, sample, evaluator group, owner, and review date. Defer low-impact variance with a trigger. If the unit or evidence view is unstable, prioritize defining it before interpreting agreement carefully.
What to carry forward
The priority output is an agreement queue tied to decision risk, dimension leverage, evaluator reach, and evidence stability. Start where disagreement could reverse a result. Preserve lower-impact variance as context instead of folding it into an average before making a release or model choice from the scores under review.
Technical background: Google DeepMind evaluation research.
Keep the decision with the work.
Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.
Create an account See the product workflow