Aglet

Triage Agent Benchmark Contamination Risk

A benchmark result can look strong because the item or answer pattern was already exposed. Triage contamination by tracing the item lifecycle and separating known exposure from plausible risk. Keep evaluation behavior and dataset hygiene as distinct questions for this review.

Establish what is happening

  1. Define the exposure boundary

    Record item source, creation date, publication, prompt text, references, retrieval corpus, tuning data, and who could access the benchmark. State which exposure types invalidate, flag, or merely qualify a result.

  2. Map contamination signals

    Inspect memorized wording, unusually high performance on public items, answer leakage, duplicated passages, benchmark-specific formatting, and changes after a hidden split. Mark observed evidence, hypothesis, and unknown separately for follow-up.

  3. Bound the score use

    State which items, model versions, runs, and claims the contamination review covers. Define quarantine or rerun criteria. Do not call a benchmark invalid from a suspicion without preserving the evidence and alternate comparison.

What to carry forward

The triage output is a contamination register with item provenance, exposure paths, signals, scope, and score treatment. Stop when reviewers know which evidence is observed. Keep a result qualified when training or retrieval history cannot be established before using it in a published comparison or decision.

Technical background: Google DeepMind evaluation research.

Keep the decision with the work.

Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.

Create an account See the product workflow