From software problem to a clear next step.
Choose a topic and what you need to do next. Add a search phrase to find a specific problem.
Browse 20 topics
Explore the playbooks
1500 guides · Page 21 of 60
-
Triage an Agent Evaluation Dataset
Scope an agent evaluation dataset by task goal, input context, outcome definition, variation, provenance, and the failures the sample should reveal.
-
Prioritize Cases for an Agent Evaluation Dataset
Choose dataset additions by failure consequence, coverage gap, task diversity, and the learning value of a case before spending effort on more examples.
-
Investigate Whether an Agent Dataset Reflects Real Tasks
Build a reproducible dataset brief by comparing evaluation cases with task variation, observed traces, label decisions, and failure modes the benchmark should expose.
-
Verify an Agent Evaluation Dataset Before Scoring
Check that an evaluation dataset has a clear task contract, useful variation, traceable labels, protected holdouts, and an outcome definition suited to the behavior under review.
-
Turn Dataset Reviews Into Better Agent Evaluation
Capture what dataset design exposed, improve case and label practice, and set a signal for detecting easy-example bias before a benchmark result is used.
-
Triage a Rubric for Agent Evaluation
Scope an agent evaluation rubric by task outcome, dimensions, score anchors, critical failures, evaluator instructions, and the decisions the score should support.
-
Prioritize Dimensions in an Agent Evaluation Rubric
Rank rubric dimensions by user consequence, decision importance, evaluator disagreement, and the value of clarifying a dimension before it influences a system comparison.
-
Investigate Whether an Agent Rubric Measures the Task
Build a reproducible rubric brief by comparing score dimensions with task requirements, anchor examples, evaluator reasoning, and failures the current rubric may hide.
-
Verify an Agent Evaluation Rubric Before Use
Check that a rubric defines task-specific dimensions, observable anchors, critical failures, evaluator instructions, and a score boundary before it drives agent comparisons.
-
Turn Rubric Reviews Into Better Agent Evaluation
Record how a rubric hid or exposed an agent failure, improve anchor and calibration practice, and set a signal for dimensions that collapse important distinctions.
-
Triage a Change in Agent Evaluation Scores
Scope an evaluation-score change by model version, dataset slice, evaluator, prompt, task distribution, run context, and the decision the movement may affect.
-
Prioritize Which Evaluation Drift to Resolve
Rank drift investigations by decision consequence, size and direction of the movement, affected slices, and the value of restoring a trustworthy comparison first.
-
Investigate Whether an Agent Score Drifted
Build a reproducible drift brief by rerunning matched slices, comparing evaluator and data versions, and separating model behavior from changes in the evaluation system.
-
Verify an Agent Evaluation Comparison Is Stable
Check that a score comparison uses matched data, evaluator, prompts, configuration, and slice definitions, and that meaningful drift is visible before the result guides a decision.
-
Turn Evaluation Drift Findings Into Better Monitoring
Capture how score movement was misread or correctly attributed, improve version and slice tracking, and set a signal for detecting measurement drift early.
-
Triage Agent Tool Selection Behavior
Scope tool-selection evaluation by task preconditions, available tools, argument rules, dependency order, expected state, and the failure the score should reveal.
-
Prioritize Tool-Use Cases for Agent Evaluation
Rank tool-selection cases by user or system consequence, likelihood of confusion, dependency complexity, and the value of catching an incorrect path before it reaches a final answer.
-
Investigate an Agent Tool-Selection Failure
Build a reproducible tool-use brief by comparing the chosen call path with valid alternatives, argument requirements, preconditions, order, and resulting state.
-
Verify Agent Tool Selection and Ordering
Check that tool calls meet schema, precondition, argument, dependency, ordering, and resulting-state requirements across the paths the task is meant to evaluate.
-
Turn Tool-Use Evaluations Into Better Agent Checks
Capture what tool-use scoring revealed, improve trace and path coverage, and set a signal for final answers that conceal invalid intermediate actions.
-
Triage Grounding in Agent Answers
Scope grounding evaluation by supplied context, answer claims, request fit, omissions, contradictions, abstention cases, and the evidence a reviewer can inspect.
-
Prioritize Grounding Cases for Agent Evaluation
Rank grounding cases by consequence of unsupported claims, context difficulty, user reliance, and the value of exposing omissions or contradictions before an answer is trusted.
-
Investigate Unsupported Claims in an Agent Answer
Build a reproducible grounding brief by extracting material claims, matching them to supplied context, and testing whether omissions, contradictions, or retrieval gaps explain the result.
-
Verify Grounded Agent Answers Against Context
Check that material answer claims are supported by the allowed context, that omissions and contradictions are visible, and that no-answer cases are judged fairly.
-
Turn Grounding Reviews Into Better Agent Evaluation
Capture how an answer exceeded or respected its context, improve claim-level review, and set a signal for unsupported specificity in future evaluations.