Aglet

Turn Rubric Reviews Into Better Agent Evaluation

A rubric review reveals which words invite inconsistent scoring and which examples make a difficult judgment concrete. Preserve the failed anchor, evaluator reasoning, and decision impact. Use that case to improve future rubric releases without pretending one scoring scheme fits every task.

Keep the lesson for the next incident

  1. Save the rubric case

    Keep the task, dimension, anchor version, evaluator rationales, counterexample, and changed score together. Include a case where a polished answer passed despite a critical path failure. That evidence shows why the dimension mattered.

  2. Improve calibration

    Add borderline and critical-failure anchors, independent scoring, and adjudication notes to rubric maintenance. Require a task-specific review when the workflow changes. Explain how the practice addresses the ambiguity found in this evaluation.

  3. Watch for score collapse

    Set a signal such as nearly identical totals for different failure patterns, repeated use of overall helpfulness, or high scores beside known task errors. Assign an owner to sample judgments and revisit the rubric when the signal appears.

What to carry forward

The learning record should connect the hidden distinction to the revised rubric practice and recurrence signal. State which dimension or anchor changed. Keep the lesson tied to the task and trace view rather than turning it into a universal judging rule.

Technical background: Agent evaluation research on arXiv.

Keep the decision with the work.

Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.

Create an account See the product workflow