Aglet

Investigate an Agent Cost-Quality Tradeoff

Cost-quality surprises usually come from mixed workloads or hidden work. Reconstruct each path at task level, separate quality dimensions, and attribute cost to calls, tokens, review, and retries. Preserve outliers so an average does not hide a poor frontier point.

Build a useful investigation brief

  1. Assemble matched runs

    Collect task slice, workload, model, prompt, tools, outputs, quality scores, latency, tokens, calls, retries, review effort, and failure costs. Record sampling and missing components. Keep useful outcome definition beside raw metrics.

  2. Split the tradeoff

    Compare quality by task and dimension against cost components. Inspect high-cost failures, cheap misses, retries, and long-context outliers. Separate workload composition, agent behavior, evaluator variance, and accounting gaps as causes.

  3. Test candidate paths

    Rerun matched slices with one model, prompt, tool, or review condition changed. Compare confidence intervals or stated variance. End with a supported frontier, a workload-specific choice, or uncertainty when quality and cost cannot be aligned.

What to carry forward

The investigation brief should link matched workload, quality slices, cost ledger, latency, retries, and candidate paths. End with an attributed tradeoff or bounded uncertainty. Do not call a cheaper path equivalent when its quality floor or useful outcome differs for the tested workload slices.

Technical background: Agent evaluation research on arXiv.

Keep the decision with the work.

Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.

Create an account See the product workflow