Aglet

Triage an Agent Evaluation Dataset

A benchmark can look rigorous while representing only the easiest form of a task. Triage the dataset before scoring by naming the user goal, allowed tools, expected state changes, and evidence needed to judge success. Make gaps and holdout boundaries visible before a result is trusted.

Establish what is happening

  1. Define the task contract

    Write the task goal, input schema, available context, allowed actions, expected output or state, and stopping condition. Separate answer quality from tool or trajectory requirements. Record cases where several valid paths should receive credit.

  2. Map useful variation

    List common, boundary, ambiguous, incomplete, and recovery cases. Record task source, version, label method, and missing context. Compare the sample with production-like work or carefully justified proxies without claiming the benchmark is representative by default.

  3. Set dataset boundaries

    Define development, calibration, and held-out splits, plus exclusion and redaction rules. State which failure modes the set cannot detect. Assign ownership for labels and revisions so a growing benchmark does not silently change its meaning.

What to carry forward

The triage brief should name the task contract, variation, provenance, split policy, and blind spots. Stop when another evaluator can see what the dataset can measure. Keep any claim provisional when the sample lacks important task or user contexts.

Technical background: Agent evaluation research on arXiv.

Keep the decision with the work.

Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.

Create an account See the product workflow