Aglet

Turn Evaluation Drift Findings Into Better Monitoring

A drift investigation teaches the team which metadata makes evaluation results interpretable. Preserve the missing version, stable comparison, and decision impact. Turn that case into a monitoring habit that flags changed measurement conditions before a score becomes a story about the agent.

Keep the lesson for the next incident

  1. Save the drift case

    Keep prior and current run metadata, slice scores, evaluator or label changes, example traces, and final interpretation together. Include a movement caused by the evaluation system rather than the agent. That contrast helps reviewers read future trends.

  2. Improve run records

    Require pinned dataset, evaluator, prompt, scoring code, model, task mix, and stable-slice metadata for each comparison. Add a change log and an inconclusive status. Explain how these fields address the ambiguity found in this investigation.

  3. Watch for metadata gaps

    Set a signal such as scores without dataset versions, changing slice counts, evaluator updates without recalibration, or sudden movement in one task family. Assign an owner to review the signal and revisit monitoring when it appears.

What to carry forward

The learning record should connect the misattributed movement to the new tracking practice and recurrence signal. State which metadata or slice check changed. Keep the lesson tied to comparison integrity rather than treating every score change as drift.

Technical background: Google DeepMind evaluation research.

Keep the decision with the work.

Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.

Create an account See the product workflow