Aglet

Investigate Whether an Agent Rubric Measures the Task

Rubric investigation asks whether a score reflects the behavior that matters or merely fluent presentation. Map each dimension to the task contract, inspect examples at the boundary, and compare evaluator interpretations. Preserve disagreements as design evidence instead of averaging them away.

Build a useful investigation brief

  1. Map rubric to task

    For each required outcome, identify the dimension and observable evidence that represents it. Mark requirements with no dimension and dimensions with no task consequence. Include tool, state, grounding, and recovery conditions when the task requires them.

  2. Review anchor decisions

    Ask evaluators to score clear passes, clear failures, and near misses without coaching. Record evidence cited, score rationale, and dimension overlap. Look for a high score in one dimension masking a critical failure elsewhere.

  3. Test a rubric change

    Revise one definition or anchor and rescore a matched sample. Compare agreement, explanations, and missed failures. State whether the change improves task validity or only makes a preferred answer easier to reward.

What to carry forward

The investigation brief should connect task requirements to rubric dimensions, anchor judgments, disagreement, and the tested revision. End with a clearer rubric or evidence request. Do not treat evaluator confidence as proof that the rubric measures the intended behavior.

Technical background: Agent evaluation research on arXiv.

Keep the decision with the work.

Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.

Create an account See the product workflow