Aglet

Triage Agent Instruction Following

An answer can sound helpful while quietly violating one decisive instruction. Triage the request into testable obligations before scoring. Separate hard constraints from preferences, and include conflicts, missing details, and valid concise responses that expose ambiguity and priority before scoring the request.

Establish what is happening

  1. Extract instruction units

    List required actions, outputs, formats, limits, exclusions, order, conditional clauses, and success state. Mark priority and dependency. Keep inferred preferences separate from words the request actually states for scoring purposes.

  2. Map request variants

    Include complete instructions, conflicting constraints, optional detail, ambiguous wording, format pressure, missing inputs, and a condition that changes the path. Record expected behavior for clarification, partial completion, and refusal for independent reviewers.

  3. Bound the score

    State whether the review measures compliance, task outcome, format, prioritization, or all four. Define severity for each missed instruction. Do not treat extra helpful text as a failure unless it violates a stated boundary.

What to carry forward

The triage output is an instruction contract with obligations, priorities, conditions, formats, exclusions, and severity rules. Stop when reviewers can identify a missed requirement. Keep inferred preferences provisional until the request or evidence supports them. Record the unresolved requirement before accepting the next output.

Technical background: OpenAI evaluation guidance.

Keep the decision with the work.

Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.

Create an account See the product workflow