Project analysis
Restoring service includes the work that accumulated
GitHub’s October 2018 incident provides a starting point for thinking about data integrity, queued work, and credible recovery criteria.
An analysis of a publicly documented project. Sources are linked in the article.
The published account
GitHub’s October 2018 incident report describes a network interruption followed by database failover into a topology its applications could not adequately support. Writes existed in different locations, complicating a safe return. The response prioritised data integrity, restored a workable topology, and processed accumulated work. The report also discusses recovery estimates that did not account for changing load. GitHub’s October 21 post-incident analysis.
This is a reading of GitHub’s public account, not a claim of involvement. The following questions are our interpretation and a proposed exercise for other teams.
Define what recovered means
For a service that accepts work, a successful response from the application is only one part of recovery. People may still be waiting for something the system accepted earlier. They need a way to find out what happened to it.
Imagine an appointment service recovering from an interruption. New bookings work again, but some confirmations remain queued. A person who received no confirmation may have tried twice. The operational task now includes finding incomplete requests, identifying duplicates, and deciding which messages remain appropriate to send.
Start the recovery plan from those user-visible states. An engineer, an operations colleague, and someone responsible for the customer experience may each notice a different part of the unfinished work. Write the states down before assigning a universal green status to the service.
Give accumulated work an operating policy
For one queue or deferred workflow, agree how to distinguish new work from work accepted before the interruption. Decide which operations can be repeated safely and which need a check against an existing result. Where repetition could affect a person twice, make the detection and resolution visible.
Specify what happens to work that has become stale. An expired invitation and an unfulfilled order should not acquire the same policy merely because both are messages. The domain matters. Include someone who understands the consequence of delivering late, delivering twice, or never delivering.
Give reconciliation an owner and a completion condition. “The queue is empty” is useful only if the team understands why it became empty. It does not explain whether records were processed, rejected, expired, or removed by an operator.
Communicate observations and uncertainty separately
A useful recovery update can describe what works now, what remains affected, and what evidence will support the next update. When an estimate depends on an assumption, expose that assumption. A confidence interval with no basis is still guesswork wearing a more formal shirt.
During a rehearsal, ask the team to write a short update using only the evidence currently available. Have a colleague outside the incident response read it. Can they tell whether their own task is safe to attempt? Can they distinguish a restored function from work that is still being reconciled?
The resulting wording belongs with the runbook, alongside the technical steps. A recovery procedure is more complete when it helps the crew explain the system’s condition to the people relying on it.
Choose one accepted-but-unfinished workflow this week. Document how the crew would locate it, resolve it, and tell its owner what happened. That is a bounded improvement you can make before the next interruption.