Aglet

Turn Adversarial Findings Into Better Agent Evaluation

Adversarial cases teach which pressure conditions reveal a meaningful gap. Preserve the clean control, stressed output, target behavior, and boundary result. Turn the case into a maintained slice with plausible conditions and clear interpretation limits for future comparisons and releases.

Keep the lesson for the next incident

  1. Save the adversarial case

    Store target, task contract, control, pressure construction, context, output, trace, evaluator, score, variants, and interpretation. Include a case that appeared surprising but failed to generalize, so novelty stays separate from evidence.

  2. Improve slice practice

    Add clean controls, pressure levels, item provenance, expected responses, recovery variants, and scope rules to evaluation plans. Explain how the change addresses the observed false positive or missed susceptibility in the controlled variant.

  3. Watch for pressure drift

    Set a signal such as failures only on one item, control cases failing, changing prompt tricks, unsupported generalization, or pressure variants no longer matching user conditions. Assign an owner to sample slices and revise them when it appears.

What to carry forward

The learning record should connect the stressed behavior to revised controls, variants, or scope practice and recurrence signal. State which pressure boundary became clear. Keep the lesson tied to plausible task conditions rather than treating adversarial novelty as capability evidence without matched controls.

Technical background: Google DeepMind evaluation research.

Keep the decision with the work.

Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.

Create an account See the product workflow