Aglet

Verify an Agent Evaluation Comparison Is Stable

Verification tests the comparison contract rather than the agent claim itself. Review versions, stable slices, labels, and score decomposition. A verified comparison can reveal a real movement or an inconclusive measurement; both outcomes are more useful than a false trend during release review.

Check whether the outcome improved

  1. Set comparison checks

    Require model and evaluator versions, prompt, dataset snapshot, task mix, scoring code, configuration, slice definitions, and run date. Add a stable core slice and a record of deviations. Define the tolerance or review rule for movement.

  2. Review matched results

    Compare current and prior scores by slice, inspect examples at material shifts, and check label or evaluator changes. Record whether the agent output, task distribution, or scoring path changed. Preserve a no-change case for calibration.

  3. Approve the claim boundary

    State whether the evidence supports improvement, regression, drift, or inconclusive movement. Record follow-up owner and trigger. Keep the historical comparison usable only within the versions and slices actually checked for this review.

What to carry forward

The verification record should show matched inputs, stable slice results, version changes, examples, and claim boundary. Use the comparison only for its tested contract. Reopen it when data, evaluator, prompt, or task distribution changes. Preserve its matched baseline for comparison.

Technical background: Google DeepMind evaluation research.

Keep the decision with the work.

Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.

Create an account See the product workflow