Decide where the work belongs
Rank consequence
Describe what happens when each dimension fails: wrong state, lost context, unsupported claim, unnecessary escalation, or wasted tool work. Identify the task and user outcome affected. Give critical constraints their own dimension when averaging would hide them.
Review ambiguity and effort
Compare anchor clarity, evaluator burden, evidence availability, and overlap with other dimensions. Prefer a smaller rubric that produces stable judgments over a long list of vague qualities. Record the examples needed to resolve the largest ambiguity.
Choose a rubric queue
Select dimensions, anchor revisions, calibration cases, owners, and review date. Defer refinements that cannot change a decision. If a dimension depends on trajectory details, prioritize making those details visible before scoring it.
What to carry forward
The priority output is a rubric queue showing consequence, ambiguity, evaluator effort, and decision value. Clarify the dimension most likely to reverse a comparison first. Keep lower-priority qualities separate rather than hiding them inside a broad aggregate.
Technical background: Agent evaluation research on arXiv.
Keep the decision with the work.
Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.
Create an account See the product workflow