Aglet

Investigate Possible Benchmark Contamination

Contamination investigation needs a chain of evidence from item creation to model access. Compare performance on exposed and controlled tasks, inspect retrieval and tuning records, and avoid using a surprising score as proof. Preserve uncertainty where history cannot be recovered.

Build a useful investigation brief

  1. Trace item history

    Collect source, authoring date, publication, prompt variants, reference answers, retrieval corpus, tuning inputs, benchmark access, and model versions. Record missing provenance. Keep the original item unchanged for audit and create separate test copies.

  2. Compare controlled slices

    Evaluate exposed, paraphrased, hidden, newly authored, and task-matched items where possible. Inspect answer overlap and unusual formatting responses. Separate item familiarity, retrieval leakage, evaluator artifacts, and actual task skill as competing explanations.

  3. Test the claim boundary

    Rerun with a fresh holdout or excluded corpus and compare score movement. Review whether the model changed or only item exposure changed. State the smallest conclusion supported and whether the original result should be qualified, replaced, or retained.

What to carry forward

The investigation brief should preserve item provenance, access history, slice comparison, overlap evidence, controlled rerun, and claim boundary. End with an attributed exposure risk or bounded uncertainty. Do not erase the original result; label how it may be used.

Technical background: Google DeepMind evaluation research.

Keep the decision with the work.

Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.

Create an account See the product workflow