Check whether the outcome improved
Set gate checks
Require baseline versions, task mix, protected slices, dimensions, scoring, evaluator, threshold, variance, raw outputs, and outcomes. Define borderline, missing-evidence, and infrastructure-failure handling plus review owner and exception path for each release candidate.
Review gate outcomes
Apply checks to pass, fail, borderline, slice regression, evaluator change, and infrastructure-fault cases. Inspect aggregate scores beside protected slices. Record whether each outcome leads to a defined action and owner.
Approve release scope
State which agent changes, slices, decisions, and evidence the gate supports. Record reviewer, owner, date, and retest trigger. Mark it inconclusive when baseline, variance, or protected-slice evidence is missing or unclear.
What to carry forward
The verification record should show baseline, slices, dimensions, threshold, variance, outcomes, and review path. Approve only the gate behavior observed. Reopen it when model, prompt, data, evaluator, scoring, task mix, or release policy changes. Record the protected slice, reviewer, and next action before accepting a pass or exception.
Technical background: OpenAI evaluation guidance.
Keep the decision with the work.
Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.
Create an account See the product workflow