Aglet

Verify Human Review Sampling for Agent Outputs

Verification tests whether human review can support the conclusion attached to it. Inspect representative and targeted samples separately, then compare the frame with reviewed cases. A passing sample makes selection and uncertainty visible before labels are generalized to a population.

Check whether the outcome improved

  1. Set sampling checks

    Require population, window, output unit, frame, selection method, strata, exclusions, reviewer context, label rules, nonresponse, and inference goal. Define when a sample is exploratory, representative, or risk-focused for each question.

  2. Review sampled coverage

    Apply checks to random, stratified, boundary, rare-risk, disagreement, low-confidence, and recent-change samples. Compare sample and population metadata. Record whether labels and reviewer context are clearly sufficient for the stated question.

  3. Approve the inference

    State which population, slices, labels, and decisions the sample supports. Record owner, reviewer, date, and retest trigger. Mark results exploratory or inconclusive when the frame or selection probabilities are missing.

What to carry forward

The verification record should show population, frame, selection, strata, exclusions, reviewer context, labels, and inference limits. Approve only the conclusion the sample supports. Reopen it when population, task mix, model, selection, reviewer, or question changes. Record the sampling owner and selection rationale before treating a targeted finding as prevalence data.

Technical background: Google DeepMind evaluation research.

Keep the decision with the work.

Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.

Create an account See the product workflow