Build a useful investigation brief
Trace item history
Collect source, authoring date, publication, prompt variants, reference answers, retrieval corpus, tuning inputs, benchmark access, and model versions. Record missing provenance. Keep the original item unchanged for audit and create separate test copies.
Compare controlled slices
Evaluate exposed, paraphrased, hidden, newly authored, and task-matched items where possible. Inspect answer overlap and unusual formatting responses. Separate item familiarity, retrieval leakage, evaluator artifacts, and actual task skill as competing explanations.
Test the claim boundary
Rerun with a fresh holdout or excluded corpus and compare score movement. Review whether the model changed or only item exposure changed. State the smallest conclusion supported and whether the original result should be qualified, replaced, or retained.
What to carry forward
The investigation brief should preserve item provenance, access history, slice comparison, overlap evidence, controlled rerun, and claim boundary. End with an attributed exposure risk or bounded uncertainty. Do not erase the original result; label how it may be used.
Technical background: Google DeepMind evaluation research.
Keep the decision with the work.
Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.
Create an account See the product workflow