Aglet

Prioritize Dimensions in an Agent Evaluation Rubric

Rubric design often starts with too many desirable qualities and no clear tradeoff. Prioritize dimensions that can change the decision or expose a severe failure. Keep secondary qualities visible, but do not let them dilute a critical task condition during review.

Decide where the work belongs

  1. Rank consequence

    Describe what happens when each dimension fails: wrong state, lost context, unsupported claim, unnecessary escalation, or wasted tool work. Identify the task and user outcome affected. Give critical constraints their own dimension when averaging would hide them.

  2. Review ambiguity and effort

    Compare anchor clarity, evaluator burden, evidence availability, and overlap with other dimensions. Prefer a smaller rubric that produces stable judgments over a long list of vague qualities. Record the examples needed to resolve the largest ambiguity.

  3. Choose a rubric queue

    Select dimensions, anchor revisions, calibration cases, owners, and review date. Defer refinements that cannot change a decision. If a dimension depends on trajectory details, prioritize making those details visible before scoring it.

What to carry forward

The priority output is a rubric queue showing consequence, ambiguity, evaluator effort, and decision value. Clarify the dimension most likely to reverse a comparison first. Keep lower-priority qualities separate rather than hiding them inside a broad aggregate.

Technical background: Agent evaluation research on arXiv.

Keep the decision with the work.

Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.

Create an account See the product workflow