Decide where the work belongs
Rank decision impact
Describe which decision relies on each score movement and what a false conclusion could cost. Give weight to critical task slices, safety-relevant behavior, or a model choice that would be hard to reverse.
Compare restoration paths
Consider rerunning a stable core set, pinning evaluator and data versions, auditing labels, or splitting the score by task. Compare time, evidence quality, and whether each path separates agent behavior from measurement change.
Choose a drift queue
Select the comparison, owner, evidence request, and review date. Defer movements that cannot affect a current decision with a stated trigger. Record when an aggregate score must be withheld from release discussion.
What to carry forward
The priority output is a drift queue tied to decision impact, affected slices, restoration path, owner, and review trigger. Restore the comparison most likely to change a decision. Keep unresolved movements visible rather than quietly folding them into a new baseline.
Technical background: Google DeepMind evaluation research.
Keep the decision with the work.
Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.
Create an account See the product workflow