Aglet

Verify recovery after a deployment rollback event

Verification should answer both questions behind rollback confidence: which version serves users and whether the original behavior recovered. Reuse the initial path and comparison window, including long-lived workers or delayed jobs when relevant.

Check whether the outcome improved

  1. Confirm serving identity

    Record current deployment, environment, service, route, worker, queue, and immutable version evidence for the affected path. Keep any participant outside the observation boundary listed as uncovered.

  2. Recheck the symptom

    Compare errors, successful completions, latency, retries, and downstream effects with the original failure and a known-good interval. A lower count without comparable traffic should remain qualified.

  3. Cover the recovery window

    Observe through the longest meaningful delay from triage, such as cache expiry or scheduled processing. Apply the decision checkpoint and leave a dated follow-up if any evidence is incomplete.

What to carry forward

Verification should say recovered, contained, still failing, or inconclusive, with serving and behavior evidence. Close only when the original path and delay window are covered; a rollback record alone does not prove recovery. Confirm the served version and affected behavior independently after the rollback checkpoint, then compare errors and critical paths with the pre-rollback baseline before closing.

Keep the decision with the work.

Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.

Create an account See the product workflow