Aglet

Investigate Whether an Agent Score Drifted

Drift investigation requires a controlled comparison before interpretation. Reconstruct the prior run, pin the relevant inputs, and break the score into slices. Check evaluator and label changes as carefully as model changes, because a different measurement can create the same headline movement.

Build a useful investigation brief

  1. Recreate the prior view

    Recover model identifier, prompt, evaluator, dataset snapshot, scoring code, task mix, and configuration. Compare hashes or version records where available. Record missing fields and do not substitute current defaults without labeling the deviation.

  2. Split the movement

    Rerun stable and changed slices separately. Compare task family, difficulty, input context, evaluator, and label distribution. Inspect examples at the largest shifts and record whether the agent output changed or only the scoring path changed.

  3. Test competing causes

    Hold one factor stable at a time when possible. Consider data composition, label revisions, evaluator calibration, prompt changes, and model behavior. State the smallest additional comparison needed when causes remain confounded or unclear.

What to carry forward

The investigation brief should preserve the prior run, matched rerun, slice movement, changed components, and competing causes. End with an attributed drift explanation or a bounded uncertainty statement. Do not report model improvement until measurement changes are separated.

Technical background: Google DeepMind evaluation research.

Keep the decision with the work.

Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.

Create an account See the product workflow