Build a useful investigation brief
Recreate the prior view
Recover model identifier, prompt, evaluator, dataset snapshot, scoring code, task mix, and configuration. Compare hashes or version records where available. Record missing fields and do not substitute current defaults without labeling the deviation.
Split the movement
Rerun stable and changed slices separately. Compare task family, difficulty, input context, evaluator, and label distribution. Inspect examples at the largest shifts and record whether the agent output changed or only the scoring path changed.
Test competing causes
Hold one factor stable at a time when possible. Consider data composition, label revisions, evaluator calibration, prompt changes, and model behavior. State the smallest additional comparison needed when causes remain confounded or unclear.
What to carry forward
The investigation brief should preserve the prior run, matched rerun, slice movement, changed components, and competing causes. End with an attributed drift explanation or a bounded uncertainty statement. Do not report model improvement until measurement changes are separated.
Technical background: Google DeepMind evaluation research.
Keep the decision with the work.
Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.
Create an account See the product workflow