Keep the lesson for the next incident
Document degraded states
Record dependency classes, local pending state, stale fallback age, retry and readback policy, freshness label, visible recovery, and side-effect limits. Define which operation classes may continue and who owns reconciliation.
Keep recovery fixtures
Retain healthy, timeout, connection, rate, permanent rejection, stale, no-fallback, recovery, read, and write cases with synthetic resources. Store expected freshness and local state. Include the original outage-shaped failure. Preserve freshness beside each fixture.
Review outage signals
Watch dependency failures, stale age, fallback use, unknown writes, reconciliation backlog, and recovery transitions by endpoint and client. Assign an owner and threshold. Close the follow-up only when essential callers exercise the degraded-state contract.
What to carry forward
Close learning with failure classes, freshness rules, fixtures, fallback and reconciliation owners, and outage signals. Keep external recovery limits explicit. The useful outcome is a truthful degraded state and safe next action, not a promise of uninterrupted dependency access. Review fallback use after recovery.
Technical background: RFC Editor reference.
Keep the decision with the work.
Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.
Create an account See the product workflow