Aglet

Turn Agreement Gaps Into Better Agent Evaluation

Agreement analysis reveals whether a score is measuring the agent or the scoring process. Preserve the disputed examples, rationales, and remedy result. Turn the lesson into a repeatable sample and review routine that keeps disagreement interpretable across future score comparisons.

Keep the lesson for the next incident

  1. Save the agreement pattern

    Store sample, evaluator pool, rubric, evidence views, labels, rationales, disagreement class, intervention, and fresh rescore. Include one stable dimension and one unstable dimension so the lesson remains diagnostic for later reviews.

  2. Improve agreement practice

    Add dimension slices, borderline examples, independent scoring, rationale review, and context checks to the evaluation plan. Explain how each addition addresses the observed source of disagreement rather than merely raising a target statistic.

  3. Watch for instability

    Set a signal such as one dimension reversing while totals hold, rising evaluator spread, missing context, or agreement falling on fresh samples. Assign an owner to inspect the signal and revisit calibration when it appears.

What to carry forward

The learning record should connect the disagreement pattern to the revised sampling or rubric practice and recurrence signal. State which slice became visible. Keep the lesson tied to the tested evaluator pool and task distribution rather than promising universal agreement indefinitely.

Technical background: Google DeepMind evaluation research.

Keep the decision with the work.

Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.

Create an account See the product workflow