Aglet

Investigate an Evaluator Agreement Gap

Agreement investigation asks why labels diverged, not only how far apart they are. Revisit the same outputs with the exact rubric and context each evaluator used. Separate a legitimate distinction from inconsistent application so the remedy matches the cause in practice.

Build a useful investigation brief

  1. Assemble independent labels

    Save evaluator, rubric version, evidence view, task, output, dimension scores, labels, rationales, and timestamps. Keep adjudicated labels separate. Sample clear, borderline, and disagreement-heavy cases for comparison on the same evaluation items.

  2. Classify the split

    Compare dimension definitions, anchors, source context, evaluator instructions, and output details. Mark disagreement from missing evidence, ambiguous rubric, evaluator drift, item difficulty, or a real behavior distinction. Record multiple causes when they interact.

  3. Test a remedy

    Apply one change such as a new anchor, shared context, evaluator pairing, or dimension split. Rescore the disputed sample and a fresh sample. State whether agreement improves, remains legitimately divided, or exposes a different measurement problem.

What to carry forward

The investigation brief should preserve independent labels, context, rationale comparison, disagreement class, remedy, and fresh rescore. End with an attributed agreement gap or bounded uncertainty. Do not use adjudication alone as evidence that independent scoring is reliable for this comparison.

Technical background: Google DeepMind evaluation research.

Keep the decision with the work.

Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.

Create an account See the product workflow