Aglet

Turn Tool-Use Evaluations Into Better Agent Checks

Tool evaluation teaches where outcome-only review loses important behavior. Preserve the invalid path, accepted alternative, state evidence, and decision impact. Turn that case into a reusable trace check that helps future evaluations distinguish correctness from luck under routine release review.

Keep the lesson for the next incident

  1. Save the tool case

    Store task fixture, tool schemas, available context, call sequence, arguments, returns, state snapshots, score, and final outcome. Include a correct final response with a faulty intermediate call. That contrast explains why trajectory evidence matters.

  2. Improve evaluation practice

    Add argument, precondition, dependency, state, alternate-path, and missing-outcome checks to the evaluation plan. Require a trace view for state-changing tasks. Explain how the practice addresses the blind spot found in this case.

  3. Watch for outcome-only review

    Set a signal such as final answers passing while calls violate schema, state snapshots missing from failures, or all paths compared with one golden sequence. Assign an owner to sample traces and revisit the checks when the signal returns.

What to carry forward

The learning record should connect the hidden call failure to the revised trace practice and recurrence signal. State which path or evidence became required. Keep the lesson tied to the task contract and available tools rather than prescribing one universal trajectory.

Technical background: Agent evaluation research on arXiv.

Keep the decision with the work.

Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.

Create an account See the product workflow