Aglet

Prioritize Adversarial Agent Evaluation Slices

Adversarial coverage should focus on plausible conditions that could change a decision or hide an unsafe path. Prioritize clear targets, meaningful controls, and slices that distinguish susceptibility from mere confusion. Keep theatrical edge cases out of the first queue until their plausibility is clear.

Decide where the work belongs

  1. Rank weakness consequence

    Describe what the target failure could change and how likely users are to encounter its condition. Give weight to state changes, misleading confidence, lost constraints, and failures hidden by a successful final format.

  2. Compare slice leverage

    Review clean controls, mild and strong pressure, conflicting context, distraction, missing evidence, and recovery. Identify which variants isolate the target behavior. Prefer a smaller diagnostic slice over many unrelated tricks.

  3. Choose an adversarial queue

    Select target, slice, control, owner, evidence request, and review date. Defer low-plausibility variants with a trigger. If the failure boundary is vague, prioritize clarifying expected behavior before adding pressure cases.

What to carry forward

The priority output is an adversarial queue tied to weakness consequence, plausibility, control quality, and diagnostic value. Start where pressure could expose a real user-facing failure. Keep novelty separate from evidence of susceptibility. Record control and target boundary before interpreting one stressed result.

Technical background: Google DeepMind evaluation research.

Keep the decision with the work.

Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.

Create an account See the product workflow