Aglet

Investigate a Human Review Sampling Gap

Sampling failures often come from who was never available to review. Investigate the frame and selection process before interpreting labels. Compare sampled outputs with the population metadata, preserve reviewer context, and record where nonresponse or exclusions directly limit the finding.

Build a useful investigation brief

  1. Reconstruct the frame

    Collect population window, task and output unit, filters, metadata, selection probabilities or reasons, reviewer assignment, context, labels, and nonresponse. Mark excluded or unavailable groups and preserve the original sample unchanged.

  2. Compare sample to population

    Inspect task mix, risk, time, model version, confidence, disagreement, and outcome distributions. Identify overrepresentation, missing strata, convenience selection, or reviewer attrition. Separate sampling bias from inconsistent human labels within the sample.

  3. Test a correction

    Draw a matched random, stratified, or missing-group sample. Compare labels and conclusion. State whether the original result holds, changes, or remains exploratory because the population frame cannot be recovered with confidence.

What to carry forward

The investigation brief should link population frame, selection, reviewed cases, missing groups, labels, correction sample, and conclusion. End with a supported inference or bounded sampling uncertainty. Do not generalize a targeted sample to the full population. Record the selection reason, missing group, and reviewer owner before stating prevalence.

Technical background: Google DeepMind evaluation research.

Keep the decision with the work.

Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.

Create an account See the product workflow