Aglet

Triage an Agent Judge Calibration Need

A score can move because evaluators interpret the rubric differently rather than because the agent changed. Triage calibration by finding disputed dimensions and missing anchors first. Separate evaluator disagreement from an agent failure, and make the review context explicit for this evaluator group.

Establish what is happening

  1. Name the judgment contract

    Record dimensions, scale, evidence allowed, unit of analysis, and what a score means. Identify whether the judge sees prompts, tools, context, traces, or only final answers. Keep unrelated quality dimensions separate.

  2. Map disagreement cases

    Collect borderline, clear-pass, clear-fail, and deliberately ambiguous examples. Record independent judgments and reasons. Mark disagreements caused by missing evidence, unclear wording, or different interpretations of the target behavior for review.

  3. Bound calibration

    Specify which evaluators, dimensions, examples, and score uses the calibration supports. Define a rule for unresolved ambiguity. Do not treat agreement on easy examples as proof that difficult cases are aligned.

What to carry forward

The triage output is a calibration contract with dimensions, evidence views, anchors, disputed cases, and scope. Stop when evaluators know what must be aligned. Keep scores provisional when the rubric or evidence view is still changing. Record the unresolved dimension and its owner before publishing a new comparison.

Technical background: Google DeepMind evaluation research.

Keep the decision with the work.

Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.

Create an account See the product workflow