Aglet

Prioritize Human Review Samples for Agent Evaluation

Review capacity should go where people can resolve the most uncertainty or catch the highest-consequence behavior. Prioritize protected slices, recent changes, disagreement, and rare but meaningful failures. Keep routine representative coverage healthy while adding targeted cases for the current decision.

Decide where the work belongs

  1. Rank review consequence

    Describe what an unseen failure or uncertain label could change. Give weight to user-visible decisions, protected behaviors, recent model changes, and slices where automated evaluation is least informative for this sample.

  2. Compare sampling value

    Review representative random, stratified, boundary, disagreement, low-confidence, rare-risk, and recent-change samples. Identify which design answers the current question. Prefer a sample with a declared frame over convenient examples for inference.

  3. Choose a review queue

    Select population, slice, selection method, reviewer, owner, and review date. Defer low-impact exploratory cases with a trigger. If the frame is incomplete, prioritize reconstructing it before inferring prevalence from labels.

What to carry forward

The priority output is a review queue tied to consequence, uncertainty, disagreement, change exposure, and coverage. Start where human evidence can change a decision. Keep targeted samples separate from prevalence estimates. Record the population frame and selection reason before using labels to describe broader agent behavior in release decisions.

Technical background: Google DeepMind evaluation research.

Keep the decision with the work.

Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.

Create an account See the product workflow