Check whether the outcome improved
Set calibration checks
Require rubric version, evidence view, dimension definitions, anchors, ambiguity rule, independent scoring, and rationale. Define acceptable disagreement and a re-review trigger. Keep evaluator identity and task context available for analysis.
Review fresh examples
Apply the calibrated rubric to clear, borderline, failure, and unseen cases. Compare scores and rationales by dimension. Inspect whether a shared score masks different reasoning or whether a difference is supported by distinct evidence.
Approve the scope
State which evaluators, examples, dimensions, and decisions the calibration supports. Record owner, date, rubric version, and retest trigger. Mark it inconclusive when the group is small or evidence views are not comparable.
What to carry forward
The verification record should show rubric, evidence views, independent scores, rationales, anchors, and scope limits. Approve calibrated use only for the tested evaluator group and task contract. Reopen it when the rubric, evaluator prompt, examples, or evidence view changes before publishing new scores.
Technical background: Google DeepMind evaluation research.
Keep the decision with the work.
Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.
Create an account See the product workflow