Establish what is happening
Define the run record
List task fixture, dataset snapshot, model, prompt, evaluator, tools, configuration, environment, seed or sampling rule, outputs, and score code. State which fields are required for repeat and which are descriptive only.
Map repeat conditions
Include deterministic, sampled, tool-dependent, time-sensitive, and failed runs. Record external state, service versions, retries, and unavailable inputs. Mark a result reproducible, repeatable with variance, or not repeatable for the current comparison.
Bound the conclusion
Specify whether the review supports exact rerun, comparable rerun, or only historical reference. Define acceptable score variation and missing-field treatment. Do not promise reproduction when a dependency or randomization rule is unknown.
What to carry forward
The triage output is a reproducibility record with inputs, versions, conditions, variance rule, and conclusion boundary. Stop when another reviewer knows what must be recovered. Keep the result provisional when execution context or scoring code is missing before the next comparison is published or used.
Technical background: OpenAI evaluation guidance.
Keep the decision with the work.
Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.
Create an account See the product workflow