Decide where the work belongs
Rank decision exposure
Describe which model or release decision uses each benchmark and what a false gain could change. Give weight to published claims, training reuse, broad retrieval access, and scores that are hard to replace later.
Compare evidence paths
Review provenance audit, hidden holdout, paraphrased items, fresh tasks, retrieval exclusion, and model-history checks. Identify which path distinguishes memorization from genuine capability. Prefer a bounded alternate measure over an unsupported accusation.
Choose a contamination queue
Select benchmark, slice, evidence request, owner, and review date. Defer low-impact hypotheses with a trigger. If provenance is missing, prioritize recording it before drawing comparisons from the score for release.
What to carry forward
The priority output is a contamination queue tied to decision exposure, visibility, reuse, evidence path, and owner. Start where leakage could change a consequential conclusion. Keep unconfirmed exposure clearly labeled while the alternate measure is built before the original score drives a release or model decision.
Technical background: Google DeepMind evaluation research.
Keep the decision with the work.
Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.
Create an account See the product workflow