Decide where the work belongs
Rank decision exposure
Describe what agent choice or budget uses each comparison and what a false equivalence could change. Give weight to recurring workload, quality gates, user-visible delay, and costs that scale with volume.
Compare frontier points
Review high-quality, low-cost, fast, slow, tool-heavy, retrying, and human-reviewed paths. Identify which measurements are comparable and which workload differences confound the result. Prefer matched slices with visible useful outcomes for users.
Choose a tradeoff queue
Select task slice, quality floor, cost fields, owner, and review date. Defer low-impact metric polish with a trigger. If cost accounting is incomplete, prioritize measuring the missing driver before choosing a frontier point.
What to carry forward
The priority output is a tradeoff queue tied to decision exposure, quality floor, cost scale, workload, and measurement completeness. Start where an apparent saving could change usefulness. Keep unmeasured costs visible in the comparison. Record the missing driver and owner before selecting a lower-cost path for release.
Technical background: Agent evaluation research on arXiv.
Keep the decision with the work.
Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.
Create an account See the product workflow