Aglet

Turn Reproducibility Gaps Into Better Agent Evaluation

Reproducibility cases teach which metadata makes an evaluation durable. Preserve the original artifact, failed recovery, controlled slice, and decision impact. Turn the lesson into a routine that makes variance and substitutions visible instead of letting a report appear more certain than its evidence.

Keep the lesson for the next incident

  1. Save the repeatability case

    Store original inputs, versions, environment, scoring path, outputs, recovery attempts, rerun differences, and final interpretation. Include a case with expected sampling variation and a case where a dependency change blocked faithful reproduction.

  2. Improve record practice

    Add identifiers, artifact links, configuration, sampling, dependency, scoring, and variance fields to evaluation plans. Require a substitution log. Explain how the revised record addresses the exact recovery gap found in this case.

  3. Watch for opaque results

    Set a signal such as scores without dataset versions, missing evaluator prompts, reruns using current defaults, unexplained variance, or raw outputs stored without scoring code. Assign an owner to inspect the signal and repair records when it appears.

What to carry forward

The learning record should connect the blocked rerun to the revised metadata or artifact practice and recurrence signal. State which condition became required. Keep the lesson tied to the tested evaluation path rather than treating every rerun difference as agent change.

Technical background: OpenAI evaluation guidance.

Keep the decision with the work.

Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.

Create an account See the product workflow