Aglet

Verify recovery after a release-linked error spike

Verification should answer whether the affected behavior returned to an acceptable state and whether the check covers the original failure. Compare the same path and population over a stated window, then look for recurrence after normal traffic resumes. A quiet graph alone is not enough.

Check whether the outcome improved

  1. Reuse the original boundary

    Check the same endpoint, job, environment, status class, and customer segment used in triage. Use a comparable pre-change or known-good window, and document any traffic or sampling difference that weakens the comparison.

  2. Look for behavior, not silence

    Confirm successful completions, latency, retries, and downstream effects alongside error counts. A lower error rate with more retries or fewer requests may hide continuing harm. Record each measure and the observation window.

  3. Set a recurrence check

    Review again after a representative load period, scheduled job, or delayed client cohort has passed. Define what threshold or example would reopen the work. Keep the verification result qualified when traffic or evidence is too thin.

What to carry forward

A useful verification result says recovered, contained, inconclusive, or still failing, with the exact observations behind it. Close or downgrade only when the original path is covered and the recurrence window has passed; otherwise leave a dated follow-up with an explicit owner.

Keep the decision with the work.

Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.

Create an account See the product workflow