Aglet

Turn Failure Labels Into Better Agent Evaluation

A failure taxonomy becomes valuable when labels change what the team does next. Preserve the evidence, neighboring categories, and remedy. Turn the case into a durable example or rule that helps reviewers separate agent behavior from tools, environment, and measurement.

Keep the lesson for the next incident

  1. Save the classification case

    Store unit, task, trace, output, tool and state evidence, evaluator context, candidate labels, final label, rationale, and outcome. Include an unknown case where the evidence correctly stopped a causal claim.

  2. Improve label practice

    Add a boundary example, multiple-label rule, unknown treatment, or required evidence field to the taxonomy. Explain how the revision addresses the misclassification or ambiguity instead of merely adding another name.

  3. Watch for taxonomy debt

    Set a signal such as most cases in “other,” labels changing after adjudication, repeated ownership disputes, missing evidence tags, or one category containing different fixes. Assign an owner to sample cases and revise the taxonomy when the signal appears.

What to carry forward

The learning record should connect the classification gap to the revised examples or evidence rule and recurrence signal. State which boundary became clearer. Keep the lesson tied to observed failure types rather than treating labels as root causes by default. across future samples.

Technical background: Agent evaluation research on arXiv.

Keep the decision with the work.

Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.

Create an account See the product workflow