Establish what is happening
Freeze endpoint state
Record endpoint, environment, status, disable or pause reason, status timestamps, receiver revision, and one safe delivery identifier. Note the last successful attempt, first failed attempt, pending count if available, and whether any effect already exists.
Trace recovery boundaries
Compare status change, configuration check, re-enable result, new delivery, durable receipt, response, and pending or replay disposition. Distinguish no delivery from delivery rejected after recovery. Keep the original attempt linked to any replay.
Bound the backlog
Group pending work by endpoint, environment, age, event identity, known effect, and recovery safety. Separate old unknowns from current failures and a duplicate already committed. Do not treat endpoint status alone as evidence that every pending event is actionable.
What to carry forward
Triage is complete when one endpoint recovery has a reproducible trigger, first failing boundary, affected backlog, and explicit unknowns. Hold bulk replay. Route status policy, configuration, receiver capacity, and event reconciliation as separate decisions with owners.
Technical background: Zendesk developer documentation.
Keep the decision with the work.
Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.
Create an account See the product workflow