Aglet

Triage a Rubric for Agent Evaluation

A single helpfulness score can hide a failed constraint, wrong tool call, or unsupported claim. Triage the rubric by identifying what success means for this task and which dimensions must remain separate. Record examples and ambiguities before evaluators begin scoring at scale. Keep the scoring audience and review context explicit.

Establish what is happening

  1. Name the dimensions

    List task correctness, constraint adherence, evidence use, action quality, recovery, or other dimensions the task actually requires. Define which failures are critical and which can coexist with a passing result.

  2. Write score anchors

    For each dimension, describe clear pass, partial, and fail examples with observable evidence. Include a near miss and a counterexample. Avoid adjectives such as good or helpful without stating what the evaluator should inspect.

  3. Bound evaluator use

    Specify input shown, trajectory or final output available, scoring order, abstention option, and escalation path for unclear cases. State which decisions the rubric can inform and what a score cannot establish.

What to carry forward

The triage output is a task-specific rubric map with dimensions, anchors, critical failures, evaluator view, and decision boundary. Stop when two evaluators can locate the same evidence. Keep scores provisional when an anchor depends on hidden context.

Technical background: Agent evaluation research on arXiv.

Keep the decision with the work.

Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.

Create an account See the product workflow