Decide where the work belongs
Rank coverage risk
Describe what a model or agent could appear to do well if a slice is absent: ignore a constraint, select the wrong tool, lose context, or fail on a boundary input. Include task consequence and who depends on the result.
Compare case sources
Review observed traces, support themes, synthetic variations, expert-authored cases, and adversarial examples. Compare provenance, labeling confidence, privacy treatment, and how closely each source matches the task contract. Do not reward volume from one easy source.
Choose a dataset queue
Select cases, slice owner, label plan, holdout placement, and review date. Defer lower-risk additions with a reason and trigger. Keep a missing-context request visible when it could change the evaluation conclusion.
What to carry forward
The priority output is a dataset queue with failure consequence, coverage gap, provenance, labeling effort, and split decision. Add the case most likely to change confidence first. Keep broad coverage claims out of the record when important task variation remains absent.
Technical background: Agent evaluation research on arXiv.
Keep the decision with the work.
Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.
Create an account See the product workflow