Aglet

Investigate Inconsistent Agent Judgments

Inconsistent judgments need the disagreement itself, not only an average. Compare what each evaluator saw, which rubric phrase they applied, and where their reasoning diverged. Preserve the task context and separate a bad example from a rubric that permits multiple readings.

Build a useful investigation brief

  1. Collect paired judgments

    Save exact rubric version, evaluator prompt or instructions, evidence view, task, output, score, rationale, and timestamp for each judgment. Pair evaluators on the same examples. Mark missing context or post-hoc explanation.

  2. Locate the interpretation split

    Compare dimension definitions, anchors, evidence thresholds, and treatment of ambiguity. Classify the split as wording, example coverage, evidence access, scale use, or an actual output difference. Avoid collapsing distinct dimensions into one disagreement.

  3. Test the correction

    Revise one anchor, instruction, or evidence view and rerun the disputed cases. Add a fresh borderline case. State whether disagreement narrows, persists for a valid reason, or reveals that the dimension should be separated.

What to carry forward

The investigation brief should link paired judgments, evidence views, interpretation split, revision, and rerun. End with a localized calibration gap or a documented valid disagreement. Do not call agreement improved if only the scoring prompt changed without new examples for comparison.

Technical background: Google DeepMind evaluation research.

Keep the decision with the work.

Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.

Create an account See the product workflow