Check whether the outcome improved
Set contamination checks
Require provenance, publication and access record, retrieval or tuning boundary, hidden holdout, fresh-item plan, overlap review, and qualification rule. Define who can inspect benchmark content and how uncertain history is recorded.
Review controlled results
Compare public, hidden, paraphrased, newly authored, and task-matched slices. Inspect answer patterns and context sources. Record whether performance differences support exposure, retrieval, item difficulty, or no clear explanation within this contamination check.
Approve score use
State which claims and model versions the benchmark supports, with any qualification. Record owner, evidence date, holdout result, and retest trigger. Mark it unfit for the intended claim when exposure cannot be bounded.
What to carry forward
The verification record should show provenance, exposure controls, holdout design, slice results, qualifications, and scope. Approve only the score use the evidence supports. Reopen the review when items, retrieval corpus, tuning data, access, or claim context changes before the next published result before reusing the score.
Technical background: Google DeepMind evaluation research.
Keep the decision with the work.
Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.
Create an account See the product workflow