Check whether the outcome improved
Set call checks
Require valid tool, schema, arguments, preconditions, dependency order, expected state, and completion condition. Mark which alternate paths are acceptable. Add a rule for missing outcomes or unavailable tools so evidence gaps are not scored as agent errors.
Review varied traces
Apply the checks to single-tool, chained, parallel, ambiguous, and recovery tasks. Inspect calls beside state snapshots and returns. Record a correct final answer that followed an invalid path and a different valid path that should pass.
Approve the score boundary
State which tool behavior the check supports and what requires trajectory or final-answer review. Record tool versions, task fixture, owner, and retest trigger. Keep a result inconclusive when state or argument evidence is incomplete.
What to carry forward
The verification record should show call rules, acceptable alternatives, varied traces, state outcomes, and evidence limits. Approve only the tool behavior observed. Retain a follow-up when hidden state or missing returns prevent a fair judgment.
Technical background: Agent evaluation research on arXiv.
Keep the decision with the work.
Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.
Create an account See the product workflow