From software problem to a clear next step.
Choose a topic and what you need to do next. Add a search phrase to find a specific problem.
Browse 20 topics
Explore the playbooks
1500 guides · Page 22 of 60
-
Triage an Agent Handoff Evaluation
Scope a handoff review by trigger, recipient, context packet, ownership signal, user expectation, and the next action the receiving person or agent must take.
-
Prioritize Agent Handoff Cases by Consequence
Rank handoff cases by consequence of delay or loss, ambiguity at the transfer point, recipient effort, and the value of detecting an unsafe continuation before users depend on it.
-
Investigate a Failed Agent Handoff
Build a reproducible handoff brief by reconstructing the trigger, packet delivered, ownership signal, recipient context, and next action.
-
Verify Agent Handoff Readiness
Check that the agent escalates under the defined condition, transfers sufficient context, identifies ownership, and gives the user a truthful next action across ordinary and ambiguous cases.
-
Turn Handoff Findings Into Better Agent Evaluation
Capture where a transfer preserved or lost task context, improve escalation and packet checks, and set a signal for recipients who lack an owner or next action.
-
Triage an Agent Trajectory Evaluation
Scope a trajectory review by task state, observations, tool calls, decisions, retries, side effects, and final outcome that a reviewer can reconstruct.
-
Prioritize Agent Trajectory Cases
Rank trajectory cases by consequence of invalid state changes, recovery difficulty, path ambiguity, and the value of exposing a failure that final-answer review would miss.
-
Investigate an Agent Trajectory Failure
Build a reproducible trajectory brief by replaying the task, locating the first divergent state, and comparing observations, calls, retries, and outcomes with an acceptable path.
-
Verify an Agent Trajectory Against State
Check that an agent trajectory respects required preconditions, observations, tool contracts, state transitions, recovery rules, and completion criteria for the task under review.
-
Turn Trajectory Results Into Better Agent Checks
Capture the first divergent state, improve transition and recovery evidence, and set a signal for correct outcomes that hide invalid or unobserved intermediate paths.
-
Triage an Agent Judge Calibration Need
Scope calibration by rubric dimensions, evaluator audience, disputed examples, score anchors, context available, and decisions affected by inconsistent judgments.
-
Prioritize Agent Judge Calibration Work
Rank calibration cases by decision consequence, disagreement frequency, ambiguity, evaluator reach, and the value of preventing a disputed score from guiding a release or quality claim.
-
Investigate Inconsistent Agent Judgments
Build a reproducible calibration brief by comparing evaluator instructions, evidence views, scores, rationales, and disputed examples for one rubric dimension.
-
Verify an Agent Judge Calibration
Check that evaluators apply the same rubric dimensions, evidence boundary, scale anchors, and ambiguity rule across clear, borderline, and failure examples.
-
Turn Judge Calibration Into Better Evaluation Practice
Capture a recurring interpretation split, improve rubric anchors and evidence views, and set a signal for evaluator disagreement before it changes an agent comparison.
-
Triage Evaluator Agreement for Agent Results
Scope agreement analysis by evaluator pool, task sample, rubric dimensions, score scale, evidence view, disagreement pattern, and decision that depends on consistent labels.
-
Prioritize Evaluator Agreement Gaps
Rank agreement gaps by decision consequence, dimension sensitivity, disagreement concentration, evaluator reach, and the value of stabilizing a score before it is used for agent comparison.
-
Investigate an Evaluator Agreement Gap
Build a reproducible agreement brief by pairing independent judgments, comparing evidence views and rationales, and classifying disagreement by rubric, item, evaluator, or agent behavior.
-
Verify Evaluator Agreement Before Comparing Agents
Check that independent evaluators apply a stable rubric and comparable evidence view, and that agreement is understood by dimension and task slice before scores guide an agent decision.
-
Turn Agreement Gaps Into Better Agent Evaluation
Capture a disagreement pattern, improve item sampling and rubric practice, and set a signal for dimension-specific evaluator instability before it distorts agent comparisons.
-
Triage Agent Benchmark Contamination Risk
Scope contamination review by item provenance, publication history, prompt exposure, retrieval sources, training or tuning inputs, benchmark access, and score decisions that may be affected.
-
Prioritize Benchmark Contamination Checks
Rank contamination checks by score consequence, exposure likelihood, benchmark visibility, item reuse, and the value of protecting a decision before a questionable result becomes a baseline.
-
Investigate Possible Benchmark Contamination
Build a reproducible contamination brief by tracing item history, comparing exposed and hidden slices, inspecting answer overlap, and testing a fresh measure.
-
Verify an Agent Benchmark Is Fit for Use
Check that benchmark items have traceable provenance, controlled exposure, an independent holdout, and a documented rule for interpreting possible leakage before scores guide a claim.
-
Turn Contamination Findings Into Better Benchmarks
Capture an exposure path or clean control, improve benchmark provenance and holdout practice, and set a signal for suspicious score gains before they become misleading agent evidence.