Check whether the outcome improved
Set grounding checks
Require claim extraction, context version, support rule, contradiction rule, omission handling, request-fit check, and abstention condition for every review. Define whether inference is allowed and how compound claims are split.
Review varied answers
Apply checks to complete, incomplete, conflicting, long-context, and unanswerable cases. Inspect claim-level exceptions alongside aggregate scores. Record a concise answer that is supported and a detailed answer that exceeds the context.
Approve the evidence boundary
State which claims and task contexts the score covers and what requires another evaluator or retrieval check. Record version, reviewer, and retest trigger. Mark results inconclusive when context or claim extraction is unreliable.
What to carry forward
The verification record should show claim rules, source matches, context slices, exceptions, and score limits. Approve grounding only within the supplied evidence boundary. Retain a follow-up when retrieval and generation failures cannot yet be separated. Ask for another review when context provenance, claim extraction, or request interpretation changes.
Technical background: Google DeepMind evaluation research.
Keep the decision with the work.
Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.
Create an account See the product workflow