Aglet

Triage Agent Evaluation Reproducibility

A score without its execution context is difficult to trust or explain. Triage reproducibility by listing the inputs and conditions that can change an outcome, then separate a repeatable evaluation from a merely archived report. Make missing metadata visible before rerunning.

Establish what is happening

  1. Define the run record

    List task fixture, dataset snapshot, model, prompt, evaluator, tools, configuration, environment, seed or sampling rule, outputs, and score code. State which fields are required for repeat and which are descriptive only.

  2. Map repeat conditions

    Include deterministic, sampled, tool-dependent, time-sensitive, and failed runs. Record external state, service versions, retries, and unavailable inputs. Mark a result reproducible, repeatable with variance, or not repeatable for the current comparison.

  3. Bound the conclusion

    Specify whether the review supports exact rerun, comparable rerun, or only historical reference. Define acceptable score variation and missing-field treatment. Do not promise reproduction when a dependency or randomization rule is unknown.

What to carry forward

The triage output is a reproducibility record with inputs, versions, conditions, variance rule, and conclusion boundary. Stop when another reviewer knows what must be recovered. Keep the result provisional when execution context or scoring code is missing before the next comparison is published or used.

Technical background: OpenAI evaluation guidance.

Keep the decision with the work.

Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.

Create an account See the product workflow