Aglet

Prioritize Benchmark Contamination Checks

Contamination review should begin with benchmarks that carry a decision and have the widest exposure surface. Prioritize public items, reused prompts, and sudden score jumps over obscure low-impact cases. Keep speculative risks visible without treating them as confirmed leakage in this benchmark.

Decide where the work belongs

  1. Rank decision exposure

    Describe which model or release decision uses each benchmark and what a false gain could change. Give weight to published claims, training reuse, broad retrieval access, and scores that are hard to replace later.

  2. Compare evidence paths

    Review provenance audit, hidden holdout, paraphrased items, fresh tasks, retrieval exclusion, and model-history checks. Identify which path distinguishes memorization from genuine capability. Prefer a bounded alternate measure over an unsupported accusation.

  3. Choose a contamination queue

    Select benchmark, slice, evidence request, owner, and review date. Defer low-impact hypotheses with a trigger. If provenance is missing, prioritize recording it before drawing comparisons from the score for release.

What to carry forward

The priority output is a contamination queue tied to decision exposure, visibility, reuse, evidence path, and owner. Start where leakage could change a consequential conclusion. Keep unconfirmed exposure clearly labeled while the alternate measure is built before the original score drives a release or model decision.

Technical background: Google DeepMind evaluation research.

Keep the decision with the work.

Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.

Create an account See the product workflow