Aglet

Investigate an Agent Failure Classification

Failure labels should explain what a reviewer can prove, not what seems plausible from the final answer. Investigate the trace, state, tools, and evaluator context. Compare adjacent categories and record why the chosen label fits or remains unknown in this case.

Build a useful investigation brief

  1. Assemble observed evidence

    Save task, output, trace, tool results, state changes, environment, rubric, and evaluator notes. Mark each fact as observed, inferred, or absent. Preserve the original failure before applying any correction or label revision.

  2. Compare neighboring labels

    Test the case against outcome, process, tool, evidence, environment, and evaluator categories. Identify the first boundary it crosses and whether multiple labels are needed. Record a plausible cause that the evidence does not establish.

  3. Check the classification

    Use a matched case or replay to test the label boundary. Change one relevant condition when possible. State whether the taxonomy distinguishes the case, needs a new category, or should keep the label unknown.

What to carry forward

The investigation brief should link observed evidence, candidate labels, boundary reasoning, replay, and uncertainty. End with a supported classification or explicit unknown. Do not turn a symptom into a causal label without evidence; keep it unknown until the evaluation record supports the cause.

Technical background: Agent evaluation research on arXiv.

Keep the decision with the work.

Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.

Create an account See the product workflow