Establish what is happening
Define the comparison unit
Record task, workload, quality dimensions, cost components, latency measure, sampling, and success boundary. State whether cost is per task, batch, or useful outcome. Keep infrastructure estimates separate from observed run measurements.
Map tradeoff cases
Include easy and difficult tasks, retries, tool-heavy work, long context, failures, human review, and partial completion. Record quality by slice and cost drivers. Mark cases where a lower cost changes the required outcome.
Bound the decision
Specify whether the review seeks minimum cost, quality threshold, latency limit, or a frontier of acceptable options. Define uncertainty and excluded costs. Do not select a path from one aggregate score when workload mix differs.
What to carry forward
The triage output is a cost-quality contract with units, quality slices, cost components, workload, uncertainty, and decision boundary. Stop when reviewers can compare useful outcomes fairly. Keep estimates provisional when important cost components are missing for this run.
Technical background: Agent evaluation research on arXiv.
Keep the decision with the work.
Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.
Create an account See the product workflow