Aglet

Prioritize Cases for an Agent Evaluation Dataset

Dataset growth should improve decision quality rather than simply increase row count. Prioritize cases that expose an important behavior, distinguish similar failure modes, or protect a critical user path. Compare label effort and maintenance with the cost of missing the case.

Decide where the work belongs

  1. Rank coverage risk

    Describe what a model or agent could appear to do well if a slice is absent: ignore a constraint, select the wrong tool, lose context, or fail on a boundary input. Include task consequence and who depends on the result.

  2. Compare case sources

    Review observed traces, support themes, synthetic variations, expert-authored cases, and adversarial examples. Compare provenance, labeling confidence, privacy treatment, and how closely each source matches the task contract. Do not reward volume from one easy source.

  3. Choose a dataset queue

    Select cases, slice owner, label plan, holdout placement, and review date. Defer lower-risk additions with a reason and trigger. Keep a missing-context request visible when it could change the evaluation conclusion.

What to carry forward

The priority output is a dataset queue with failure consequence, coverage gap, provenance, labeling effort, and split decision. Add the case most likely to change confidence first. Keep broad coverage claims out of the record when important task variation remains absent.

Technical background: Agent evaluation research on arXiv.

Keep the decision with the work.

Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.

Create an account See the product workflow