Aglet

Triage Human Review Sampling for Agents

Human review is useful only when the sample answers a clear question. Triage what population needs inspection, which cases need purposeful oversampling, and how findings will be interpreted. Keep representative coverage separate from risk-focused review for a defined evaluation question.

Establish what is happening

  1. Define the population

    Record time window, task types, output unit, filters, available metadata, and sampling frame. State whether review estimates prevalence, finds failures, calibrates evaluators, or explores a new behavior in this population.

  2. Map sampling needs

    Include random, stratified, boundary, rare-risk, disagreement, low-confidence, long-tail, and recent-change cases. Record inclusion probability or selection reason. Define reviewer context, redaction, and what counts as a usable review for the question.

  3. Bound the inference

    Specify which population, slices, labels, and decisions the sample supports. Define uncertainty and nonresponse treatment. Do not report a risk-focused sample as prevalence without a valid population frame for that claim.

What to carry forward

The triage output is a human-review sampling contract with population, frame, strata, selection reasons, reviewer context, and inference limits. Stop when reviewers know what the sample can answer. Keep convenience samples explicitly exploratory. Record excluded groups and the review owner before making a population claim from a targeted sample.

Technical background: Google DeepMind evaluation research.

Keep the decision with the work.

Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.

Create an account See the product workflow