Check whether the outcome improved
Set comparison checks
Require model and evaluator versions, prompt, dataset snapshot, task mix, scoring code, configuration, slice definitions, and run date. Add a stable core slice and a record of deviations. Define the tolerance or review rule for movement.
Review matched results
Compare current and prior scores by slice, inspect examples at material shifts, and check label or evaluator changes. Record whether the agent output, task distribution, or scoring path changed. Preserve a no-change case for calibration.
Approve the claim boundary
State whether the evidence supports improvement, regression, drift, or inconclusive movement. Record follow-up owner and trigger. Keep the historical comparison usable only within the versions and slices actually checked for this review.
What to carry forward
The verification record should show matched inputs, stable slice results, version changes, examples, and claim boundary. Use the comparison only for its tested contract. Reopen it when data, evaluator, prompt, or task distribution changes. Preserve its matched baseline for comparison.
Technical background: Google DeepMind evaluation research.
Keep the decision with the work.
Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.
Create an account See the product workflow