Establish what is happening
Define the exposure boundary
Record item source, creation date, publication, prompt text, references, retrieval corpus, tuning data, and who could access the benchmark. State which exposure types invalidate, flag, or merely qualify a result.
Map contamination signals
Inspect memorized wording, unusually high performance on public items, answer leakage, duplicated passages, benchmark-specific formatting, and changes after a hidden split. Mark observed evidence, hypothesis, and unknown separately for follow-up.
Bound the score use
State which items, model versions, runs, and claims the contamination review covers. Define quarantine or rerun criteria. Do not call a benchmark invalid from a suspicion without preserving the evidence and alternate comparison.
What to carry forward
The triage output is a contamination register with item provenance, exposure paths, signals, scope, and score treatment. Stop when reviewers know which evidence is observed. Keep a result qualified when training or retrieval history cannot be established before using it in a published comparison or decision.
Technical background: Google DeepMind evaluation research.
Keep the decision with the work.
Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.
Create an account See the product workflow