Build a useful investigation brief
Rebuild the comparison
Collect release change, baseline model, prompt, dataset, evaluator, scoring code, task mix, protected slices, scores, variance, and environment. Record missing or substituted fields. Preserve raw outputs and gate configuration together.
Split the crossing
Compare aggregate and protected slices, examples at material movement, evaluator and label changes, tool or environment errors, and score variance. Classify agent regression, measurement drift, infrastructure fault, or insufficient evidence as competing causes.
Test the gate path
Rerun stable slices and a fresh matched sample. Hold one condition stable at a time where possible. State whether the threshold crossing persists, moves, or disappears, and record the smallest follow-up needed.
What to carry forward
The investigation brief should link gate configuration, baseline, changed run, slices, variance, examples, and rerun. End with an attributed regression or bounded uncertainty. Do not override a gate silently because the aggregate score looks acceptable. Record the protected slice result and reviewer before making a release exception.
Technical background: OpenAI evaluation guidance.
Keep the decision with the work.
Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.
Create an account See the product workflow