Aglet

Turn Cost-Quality Results Into Better Agent Evaluation

Cost-quality cases reveal which metric is hiding useful work. Preserve matched tasks, quality slice, cost ledger, and decision impact. Turn the result into a repeatable comparison habit that keeps budget choices tied to outcome instead of a single average score.

Keep the lesson for the next incident

  1. Save the tradeoff case

    Store workload, candidate path, quality dimensions, outcomes, latency, tokens, calls, retries, review effort, cost, outliers, and decision. Include a low-cost path that missed a quality floor and a high-cost path that met it.

  2. Improve comparison practice

    Add matched slices, quality floors, full cost ledger, retry and review fields, outlier inspection, and workload labels to evaluation plans. Explain how each change addresses the hidden cost or mixed comparison.

  3. Watch for false savings

    Set a signal such as cost falling while retries rise, review effort missing, quality drops on hard slices, or averages improving after workload changes. Assign an owner to sample ledgers and revisit the frontier when it appears.

What to carry forward

The learning record should connect the hidden cost or quality gap to revised comparison fields and recurrence signal. State which floor or cost driver became visible. Keep the lesson tied to the tested workload rather than claiming one universal price-quality rule.

Technical background: Agent evaluation research on arXiv.

Keep the decision with the work.

Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.

Create an account See the product workflow