Aglet

Verify an Agent Benchmark Is Fit for Use

Verification protects the meaning of a benchmark result. Review item history, access boundaries, exposed and hidden slices, and score caveats. A benchmark can remain useful with a qualification when its limitations are explicit and the comparison uses an appropriate control.

Check whether the outcome improved

  1. Set contamination checks

    Require provenance, publication and access record, retrieval or tuning boundary, hidden holdout, fresh-item plan, overlap review, and qualification rule. Define who can inspect benchmark content and how uncertain history is recorded.

  2. Review controlled results

    Compare public, hidden, paraphrased, newly authored, and task-matched slices. Inspect answer patterns and context sources. Record whether performance differences support exposure, retrieval, item difficulty, or no clear explanation within this contamination check.

  3. Approve score use

    State which claims and model versions the benchmark supports, with any qualification. Record owner, evidence date, holdout result, and retest trigger. Mark it unfit for the intended claim when exposure cannot be bounded.

What to carry forward

The verification record should show provenance, exposure controls, holdout design, slice results, qualifications, and scope. Approve only the score use the evidence supports. Reopen the review when items, retrieval corpus, tuning data, access, or claim context changes before the next published result before reusing the score.

Technical background: Google DeepMind evaluation research.

Keep the decision with the work.

Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.

Create an account See the product workflow