Aglet

Verify Agent Tool Selection and Ordering

Verification asks whether the trace supports a fair judgment of tool behavior. Apply explicit rules to ordinary, ambiguous, and recovery cases, then inspect exceptions. A passing tool path should be defined by valid state and task outcome, not by one preferred sequence without state evidence.

Check whether the outcome improved

  1. Set call checks

    Require valid tool, schema, arguments, preconditions, dependency order, expected state, and completion condition. Mark which alternate paths are acceptable. Add a rule for missing outcomes or unavailable tools so evidence gaps are not scored as agent errors.

  2. Review varied traces

    Apply the checks to single-tool, chained, parallel, ambiguous, and recovery tasks. Inspect calls beside state snapshots and returns. Record a correct final answer that followed an invalid path and a different valid path that should pass.

  3. Approve the score boundary

    State which tool behavior the check supports and what requires trajectory or final-answer review. Record tool versions, task fixture, owner, and retest trigger. Keep a result inconclusive when state or argument evidence is incomplete.

What to carry forward

The verification record should show call rules, acceptable alternatives, varied traces, state outcomes, and evidence limits. Approve only the tool behavior observed. Retain a follow-up when hidden state or missing returns prevent a fair judgment.

Technical background: Agent evaluation research on arXiv.

Keep the decision with the work.

Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.

Create an account See the product workflow