Decide where the work belongs
Rank label consequence
Describe what a wrong or missing label could cause: misassigned work, repeated regressions, invalid comparison, or hidden environment fault. Give weight to labels used in gates, dashboards, or recurring failure review.
Compare diagnostic value
Review whether a category distinguishes outcome, trajectory, tool, evidence, environment, or evaluator causes. Prefer labels that lead to different investigations. Split a broad bucket only when examples show a repeatable boundary.
Choose a taxonomy queue
Select category, example set, owner, and review date. Defer low-value wording changes with a trigger. If labels conflict, prioritize paired examples and evidence rules before expanding the category list further.
What to carry forward
The priority output is a taxonomy queue tied to decision consequence, diagnostic value, disagreement, and ownership. Start where one label would send cases to different fixes. Keep rare but meaningful categories visible rather than hiding them in “other” when the next review uses labels to route distinct fixes.
Technical background: Agent evaluation research on arXiv.
Keep the decision with the work.
Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.
Create an account See the product workflow