Check whether the outcome improved
Set repeatability checks
Require dataset and task identifiers, model, prompt, evaluator, tool and environment versions, configuration, sampling rule, scoring code, outputs, and timestamps. Define exact-match, tolerance, and unreproducible outcomes for every recorded run.
Review the rerun
Ask an independent reviewer to follow the record and repeat a fixed slice. Compare inputs, traces, labels, scores, and expected variance. Record every manual substitution or dependency mismatch instead of treating it as invisible setup.
Approve the scope
State whether the record supports exact reproduction, comparable measurement, or historical reference only. Record owner, date, artifacts, and retest trigger. Mark the result inconclusive when a material condition or scoring path cannot be recovered.
What to carry forward
The verification record should show run inputs, versions, conditions, rerun, variance, substitutions, and scope. Approve only the reproducibility claim the evidence supports. Reopen it when data, evaluator, model, tools, environment, or scoring code changes before another result is published or compared publicly.
Technical background: OpenAI evaluation guidance.
Keep the decision with the work.
Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.
Create an account See the product workflow