Aglet

Triage Evaluator Agreement for Agent Results

An overall agreement number can hide one dimension where evaluators reverse their judgments. Triage by splitting scores and rationales by task and dimension. Identify whether disagreement reflects rubric ambiguity, missing evidence, or a meaningful agent behavior difference in this sample.

Establish what is happening

  1. Define the agreement unit

    Record evaluator population, item sample, rubric version, dimensions, scale, evidence view, and whether agreement means exact score, acceptable range, or shared label. Keep independent judgments separate from adjudicated outcomes for analysis.

  2. Map disagreement shape

    Inspect confusion pairs, score spread, dimension reversals, task slices, and rationales. Mark clusters caused by borderline wording, context gaps, evaluator expertise, or real output variation. Preserve the items that reveal the pattern.

  3. Bound the conclusion

    State which evaluator group, sample, dimension, and decision the agreement evidence covers. Define when low agreement requires calibration or exclusion. Do not treat high aggregate agreement as proof every dimension is reliable.

What to carry forward

The triage output is an agreement brief with unit, evaluator pool, disagreement slices, causes to test, and conclusion boundary. Stop when a reviewer can see where agreement holds and fails. Keep use provisional where evidence views differ before releasing a new comparison result. from that sample.

Technical background: Google DeepMind evaluation research.

Keep the decision with the work.

Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.

Create an account See the product workflow