Aglet

Prioritize Agent Evaluation Reproducibility Work

Reproducibility work should begin where an opaque result can change an important choice. Prioritize missing model, dataset, evaluator, tool, and scoring versions over cosmetic report detail. Keep expensive environmental differences visible without assuming every variation invalidates a run or comparison.

Decide where the work belongs

  1. Rank decision exposure

    Describe which comparison or gate uses each result and what a changed rerun could alter. Give weight to high-impact claims, volatile dependencies, and evaluations that will be cited after the original environment disappears.

  2. Compare recovery paths

    Review pinning inputs, exporting artifacts, recording environment state, rerunning a fixed slice, or documenting variance. Identify which path restores interpretability fastest. Prefer a smaller controlled rerun over an unbounded attempt to recreate every dependency.

  3. Choose a repeatability queue

    Select result, missing field, owner, recovery path, and review date. Defer low-impact metadata gaps with a trigger. If scoring code cannot be recovered, prioritize preserving the raw outputs and a bounded comparison.

What to carry forward

The priority output is a reproducibility queue tied to decision exposure, dependency volatility, recovery cost, and evidence quality. Start where missing context could reverse a conclusion. Keep report polish separate from the records needed to repeat the result when another reviewer repeats the run.

Technical background: OpenAI evaluation guidance.

Keep the decision with the work.

Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.

Create an account See the product workflow