Aglet

Investigate an Adversarial Agent Failure

A stressed output can fail because the task became ambiguous or because the target behavior broke under pressure. Investigate clean controls beside adversarial variants. Preserve prompt, context, evaluator view, and recovery so the failure boundary remains clearly legible to reviewers.

Build a useful investigation brief

  1. Reconstruct the control

    Record task, target behavior, clean prompt, context, tools, expected response, agent output, trace, evaluator, and score. Preserve adversarial item construction and provenance. Mark any hidden assumption or missing evidence for reproducibility.

  2. Compare pressure variants

    Run clean, mild, strong, contradictory, distracting, incomplete, and recovery cases. Inspect whether the same target weakness appears. Separate prompt ambiguity, context conflict, tool effects, evaluator interpretation, and actual susceptibility in evidence.

  3. Test the boundary

    Change one pressure feature or task condition at a time. Add a fresh matched item. State whether the failure generalizes to the defined slice, remains an isolated example, or cannot be attributed because the control was inadequate.

What to carry forward

The investigation brief should link target, clean control, pressure variants, outputs, traces, and boundary test. End with an attributed susceptibility or bounded uncertainty. Do not call an adversarial trick a general weakness without matched evidence across task families or reviewer samples.

Technical background: Google DeepMind evaluation research.

Keep the decision with the work.

Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.

Create an account See the product workflow