Check whether the outcome improved
Set rubric checks
Require task outcome, dimension definition, score anchors, critical-failure rule, evidence view, unclear-case path, and version. Add a check that a fluent final response cannot pass when a required intermediate or constraint failed.
Run calibration cases
Score clear, borderline, and counterexample traces independently. Compare evidence cited and dimension scores. Revise wording or anchors when evaluators disagree for different reasons, not merely when their totals differ across cases.
Approve the scoring boundary
State which task versions and trace views the rubric supports. Record known blind spots, owner, and review date. Keep a manual or alternate check for failures the rubric cannot observe reliably.
What to carry forward
The verification record should show task mapping, anchor cases, evaluator agreement, critical failures, and rubric limits. Use the rubric within its tested scope. Retain a follow-up when hidden state or missing trace detail could change a score.
Technical background: Agent evaluation research on arXiv.
Keep the decision with the work.
Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.
Create an account See the product workflow