Aglet

Prioritize Which Evaluation Drift to Resolve

Drift can consume endless analysis unless the team starts with the movement that could change an important release or model decision. Prioritize changes that affect critical slices or make historical comparisons unreliable. Keep small exploratory shifts visible without treating them as urgent regressions.

Decide where the work belongs

  1. Rank decision impact

    Describe which decision relies on each score movement and what a false conclusion could cost. Give weight to critical task slices, safety-relevant behavior, or a model choice that would be hard to reverse.

  2. Compare restoration paths

    Consider rerunning a stable core set, pinning evaluator and data versions, auditing labels, or splitting the score by task. Compare time, evidence quality, and whether each path separates agent behavior from measurement change.

  3. Choose a drift queue

    Select the comparison, owner, evidence request, and review date. Defer movements that cannot affect a current decision with a stated trigger. Record when an aggregate score must be withheld from release discussion.

What to carry forward

The priority output is a drift queue tied to decision impact, affected slices, restoration path, owner, and review trigger. Restore the comparison most likely to change a decision. Keep unresolved movements visible rather than quietly folding them into a new baseline.

Technical background: Google DeepMind evaluation research.

Keep the decision with the work.

Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.

Create an account See the product workflow