Decide where the work belongs
Rank decision exposure
Describe which model, release, or workflow decision uses each judgment and what a false score could cause. Give weight to high-impact dimensions, broad evaluator coverage, and cases that are hard to revisit later.
Compare anchor value
Review clear, borderline, failure, and exception examples. Identify which examples clarify a scale boundary, evidence requirement, or allowed inference. Prefer an anchor that resolves a recurring ambiguity over several nearly identical examples.
Choose a calibration queue
Select dimension, cases, evaluators, owner, and review date. Defer low-impact wording debates with a trigger. If the evidence view differs across evaluators, prioritize a shared view before discussing score thresholds.
What to carry forward
The priority output is a calibration queue tied to decision exposure, ambiguity, disagreement, anchor value, and evaluator reach. Start where inconsistent interpretation can change a decision. Keep unresolved edge cases visible rather than averaging them away before scores enter any release comparison or report.
Technical background: Google DeepMind evaluation research.
Keep the decision with the work.
Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.
Create an account See the product workflow