Aglet

Triage a Change in Agent Evaluation Scores

A score that moves between runs does not identify its own cause. Triage drift by freezing the comparison inputs, listing changes, and separating agent behavior from benchmark or evaluator movement. Make the matched baseline visible before describing improvement or regression.

Establish what is happening

  1. Name the comparison

    Record current and prior model, prompt, evaluator, dataset version, task mix, configuration, and time window. Define the score and slice boundaries. Identify which fields must match before the runs can be compared directly.

  2. Map possible drift

    List changes in tasks, labels, traffic, tool availability, context length, evaluator prompts, and scoring code. Mark each as observed change, suspected change, or unknown. Preserve a stable core slice for a cleaner comparison.

  3. Bound the claim

    State whether the movement is compatible with agent improvement, evaluation drift, or both. Identify the next matched run or evidence request. Do not attach a causal story to an aggregate score that mixes changed inputs.

What to carry forward

The triage output is a drift register with matched fields, changed inputs, stable slices, and an evidence boundary. Stop when the comparison can be reproduced. Keep the score change provisional when a dataset or evaluator version is unknown for this comparison.

Technical background: Google DeepMind evaluation research.

Keep the decision with the work.

Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.

Create an account See the product workflow