From software problem to a clear next step.
Choose a topic and what you need to do next. Add a search phrase to find a specific problem.
Browse 20 topics
Explore the playbooks
1500 guides · Page 24 of 60
-
Triage Multi-Agent Coordination
Scope coordination evaluation by agent roles, shared task state, messages, handoffs, dependencies, conflict resolution, duplicate work, and the outcome the team must achieve.
-
Prioritize Multi-Agent Coordination Cases
Rank coordination cases by conflicting or missing work, shared-state risk, dependency complexity, recovery cost, and decision value.
-
Investigate a Multi-Agent Coordination Failure
Build a reproducible coordination brief by reconstructing roles, messages, shared state, dependencies, and the first ownership or synchronization divergence.
-
Verify Multi-Agent Coordination
Check that agents follow role boundaries, exchange sufficient context, maintain shared state, resolve conflicts, and signal ownership across serial, parallel, failed, and recovery paths.
-
Turn Coordination Findings Into Better Agent Evaluation
Capture a role, message, state, or recovery failure, improve multi-agent trace checks, and set a signal for duplicated or missing work that a final synthesis can hide.
-
Triage Agent Cost and Quality Tradeoffs
Scope cost-quality evaluation by task outcome, quality dimensions, latency, tokens, tool calls, external work, review effort, failure cost, and the decision the comparison will inform.
-
Prioritize Agent Cost-Quality Evaluations
Rank cost-quality cases by decision consequence, quality shortfall, cost exposure, workload frequency, and the value of finding a better frontier before an agent choice becomes a default.
-
Investigate an Agent Cost-Quality Tradeoff
Build a reproducible tradeoff brief by comparing matched workloads, quality slices, cost drivers, latency, retries, and useful outcomes across candidate agent paths.
-
Verify an Agent Cost-Quality Comparison
Check that agent options are compared on matched workloads with quality floors, cost components, latency, variance, and useful-outcome definitions before a tradeoff guides selection.
-
Turn Cost-Quality Results Into Better Agent Evaluation
Capture a useful frontier or hidden cost, improve matched workload and quality-floor checks, and set a signal for savings that rely on retries, review, or silent task misses.
-
Triage an Adversarial Agent Evaluation Slice
Scope an adversarial slice by target behavior, attack or pressure condition, valid task, expected evidence, failure boundary, evaluator visibility, and the decision the slice should inform.
-
Prioritize Adversarial Agent Evaluation Slices
Rank adversarial slices by target consequence, pressure plausibility, exposure, diagnostic value, and decision value.
-
Investigate an Adversarial Agent Failure
Build a reproducible adversarial brief by comparing clean and pressured cases, reconstructing the task contract, and testing whether the weakness persists across controlled variants.
-
Verify an Adversarial Evaluation Slice
Check that an adversarial slice tests a defined behavior with a clean control, pressure variants, evidence, and bounded claims.
-
Turn Adversarial Findings Into Better Agent Evaluation
Capture a defined weakness or clean control, improve adversarial slice construction, and set a signal for pressure cases that expose constraint loss or unsupported confidence.
-
Triage an Agent Evaluation Regression Gate
Scope a regression gate by baseline, task slices, quality dimensions, threshold, variance, evaluator, release decision, and the evidence required when a change crosses the gate.
-
Prioritize Agent Regression Gate Coverage
Rank gate gaps by consequence of a missed regression, protected-slice exposure, change frequency, false-block cost, and the value of making a release decision from a trustworthy baseline.
-
Investigate an Agent Regression Gate Crossing
Build a reproducible gate brief by comparing the changed run with its pinned baseline, protected slices, variance, evaluator, and infrastructure evidence.
-
Verify an Agent Regression Gate
Check that a regression gate uses a pinned baseline, protected slices, stable scoring, variance, thresholds, and a review path for borderline or missing evidence.
-
Turn Regression Gate Results Into Better Agent Evaluation
Capture a missed, noisy, or correctly detected regression, improve baseline and slice practice, and set a signal for gate crossings caused by measurement or infrastructure change.
-
Triage Human Review Sampling for Agents
Scope review sampling by output population, task slices, risk, uncertainty, evaluator disagreement, volume, reviewer expertise, and the decision the sample should inform.
-
Prioritize Human Review Samples for Agent Evaluation
Rank review samples by decision consequence, uncertainty, disagreement, change exposure, population coverage, and the value of seeing a failure before automated scores are trusted.
-
Investigate a Human Review Sampling Gap
Build a reproducible sampling brief by reconstructing the population frame, selection method, reviewed cases, labels, disagreement, and missing groups that could change the conclusion.
-
Verify Human Review Sampling for Agent Outputs
Check that review samples have a defined population frame, selection method, strata, exclusions, reviewer context, labels, and inference limits.
-
Turn Human Review Sampling Into Better Agent Evaluation
Capture missing groups, targeted samples, or reviewer disagreement; improve sampling frames and signal convenience examples presented as representative evidence.