Establish what is happening
Name the comparison
Record current and prior model, prompt, evaluator, dataset version, task mix, configuration, and time window. Define the score and slice boundaries. Identify which fields must match before the runs can be compared directly.
Map possible drift
List changes in tasks, labels, traffic, tool availability, context length, evaluator prompts, and scoring code. Mark each as observed change, suspected change, or unknown. Preserve a stable core slice for a cleaner comparison.
Bound the claim
State whether the movement is compatible with agent improvement, evaluation drift, or both. Identify the next matched run or evidence request. Do not attach a causal story to an aggregate score that mixes changed inputs.
What to carry forward
The triage output is a drift register with matched fields, changed inputs, stable slices, and an evidence boundary. Stop when the comparison can be reproduced. Keep the score change provisional when a dataset or evaluator version is unknown for this comparison.
Technical background: Google DeepMind evaluation research.
Keep the decision with the work.
Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.
Create an account See the product workflow