From software problem to a clear next step.
Choose a topic and what you need to do next. Add a search phrase to find a specific problem.
Browse 20 topics
Explore the playbooks
1500 guides · Page 23 of 60
-
Triage Agent Evaluation Reproducibility
Scope reproducibility by task inputs, model and evaluator versions, runtime configuration, tool state, scoring code, randomization, and the result another reviewer must recreate.
-
Prioritize Agent Evaluation Reproducibility Work
Rank reproducibility gaps by decision consequence, likelihood of result change, dependency volatility, rerun cost, and preserving a comparison before it becomes the baseline.
-
Investigate a Non-Reproducible Agent Evaluation
Build a reproducibility brief by reconstructing inputs and versions, rerunning a controlled slice, and separating true agent movement from changed execution conditions.
-
Verify an Agent Evaluation Can Be Repeated
Check that an evaluator can recover the task, versions, execution conditions, scoring path, artifacts, and variance rule needed to reproduce or fairly compare an agent result.
-
Turn Reproducibility Gaps Into Better Agent Evaluation
Capture which missing condition blocked a repeat, improve run records and artifact retention, and set a signal for results that cannot be rerun before they guide another agent comparison.
-
Triage an Agent Evaluation Failure Taxonomy
Scope failure classification by task outcome, trajectory state, tool behavior, evidence availability, environment condition, evaluator rule, and the decision the label will support.
-
Prioritize Agent Failure Taxonomy Gaps
Rank taxonomy gaps by decision consequence, frequency, diagnostic value, evaluator disagreement, and distinguishing a fixable agent fault from an environment or measurement problem.
-
Investigate an Agent Failure Classification
Build a classification brief by comparing observed trajectory and outcome with taxonomy definitions, neighboring labels, evidence gaps, and competing causes.
-
Verify Agent Failure Labels Are Consistent
Check that evaluators apply failure labels to the right unit, separate symptoms from causes, and preserve multi-cause evidence.
-
Turn Failure Labels Into Better Agent Evaluation
Capture a misclassified or usefully separated failure, improve taxonomy examples and evidence rules, and set a signal for broad “failed” labels that block the next investigation.
-
Triage Agent Abstention Behavior
Scope abstention by answerability, evidence sufficiency, uncertainty, request intent, clarification need, refusal boundary, and action after the agent declines or qualifies an answer.
-
Prioritize Agent Abstention Cases
Rank abstention cases by unsupported confidence, unnecessary refusal, evidence ambiguity, user reliance, and decision value.
-
Investigate an Agent Abstention Error
Build an abstention brief from available evidence and request intent; compare whether the agent guessed, refused, clarified, or helped appropriately.
-
Verify Agent Abstention Decisions
Check that an agent answers supported requests, qualifies uncertainty, asks for missing information, and declines unsupported tasks with a truthful next action.
-
Turn Abstention Findings Into Better Agent Evaluation
Capture where the agent guessed, refused unnecessarily, or gave partial help, improve answerability and evidence checks, and set a signal for confidence that exceeds available context.
-
Triage Agent Instruction Following
Scope instruction-following evaluation by explicit requirements, priorities, dependencies, format, exclusions, conditional branches, tool constraints, and required outcome.
-
Prioritize Agent Instruction Following Cases
Rank instruction cases by missed-constraint consequence, silent-failure risk, dependency complexity, visibility, and score value.
-
Investigate an Agent Instruction-Following Failure
Build a reproducible compliance brief by extracting the request contract, comparing required and actual behavior, and replaying the smallest missing or conflicting instruction.
-
Verify Agent Instruction Compliance
Check agent compliance with required actions, priorities, conditions, exclusions, formats, and completion criteria.
-
Turn Instruction Failures Into Better Agent Evaluation
Capture missed requirements or conflicts, improve obligation and branch checks, and signal outputs that violate decisive constraints.
-
Triage Retrieval Coverage for Agent Tasks
Scope retrieval coverage by task claims, required evidence, source set, ranking, omissions, distractors, provenance, and the answer or action the context is meant to support.
-
Prioritize Retrieval Coverage Cases
Rank retrieval cases by consequence of missing evidence, likelihood of silent omission, source difficulty, user reliance, and the value of fixing coverage before judging an agent answer.
-
Investigate a Retrieval Coverage Failure
Build a reproducible coverage brief by mapping required evidence to retrieved sources, testing query and ranking changes, and separating missing retrieval from ignored context.
-
Verify Retrieval Coverage Before Judging Answers
Check that retrieved context contains required evidence, provenance, freshness, omissions, and conflicts, and is scored separately from agent use.
-
Turn Retrieval Gaps Into Better Agent Evaluation
Capture a missing, noisy, or well-covered evidence path, improve retrieval slices and provenance checks, and set a signal for answers that appear grounded despite incomplete context.