Establish what is happening
Define the task contract
Write the task goal, input schema, available context, allowed actions, expected output or state, and stopping condition. Separate answer quality from tool or trajectory requirements. Record cases where several valid paths should receive credit.
Map useful variation
List common, boundary, ambiguous, incomplete, and recovery cases. Record task source, version, label method, and missing context. Compare the sample with production-like work or carefully justified proxies without claiming the benchmark is representative by default.
Set dataset boundaries
Define development, calibration, and held-out splits, plus exclusion and redaction rules. State which failure modes the set cannot detect. Assign ownership for labels and revisions so a growing benchmark does not silently change its meaning.
What to carry forward
The triage brief should name the task contract, variation, provenance, split policy, and blind spots. Stop when another evaluator can see what the dataset can measure. Keep any claim provisional when the sample lacks important task or user contexts.
Technical background: Agent evaluation research on arXiv.
Keep the decision with the work.
Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.
Create an account See the product workflow