Project analysis
A migration needs a way back at every stage
GitHub’s MySQL upgrade is a useful starting point for examining recovery, intermediate states, and the work a migration leaves behind.
An analysis of a publicly documented project. Sources are linked in the article.
The published account
In its account of moving to MySQL 8.0, GitHub describes a staged upgrade: testing application compatibility, introducing upgraded replicas, promoting a new primary, and eventually removing the older servers. The team kept recovery options available during the transition. Supporting mixed versions required changes to tooling and attention to replication compatibility. Some query problems appeared under production workloads despite passing continuous integration. GitHub’s engineering account, published in December 2023.
This is an analysis of that public account. BE A VIKING did not participate in the project. The interpretation and exercise below are ours.
Our interpretation
A migration proposal usually gives its destination plenty of attention. The intermediate states deserve their own design review. Those states are where operators must recognise what is happening, contributors must understand what they may change, and users must continue doing useful work.
Try reviewing a migration as a sequence of operating conditions. For each condition, write down which version receives new work, which version owns the authoritative records, and which tools can explain the current state. Give each condition a way to end: advance, return, or hold while a named uncertainty is investigated.
That review can change the first useful slice. A team may discover that its immediate need is a version-aware diagnostic command or a rehearsal environment. That foundation has a specific beneficiary: the next responsible change. It should not become an invitation to redesign the entire platform.
Make recovery a task someone can perform
Consider a hypothetical change to an order-processing service. Switching requests back to the earlier release may be easy. Understanding which orders the newer release accepted may be harder. A useful rehearsal includes both parts.
Ask someone who did not write the migration script to describe the recovery sequence using the available instructions. Let them identify missing access, ambiguous checks, and assumptions about the data. Keep the rehearsal bounded and use an isolated environment with representative records.
Record the point at which returning becomes a different kind of operation. Before a data conversion, it might involve changing a routing setting. Afterwards, it might require restoring records or reconciling two representations. Calling both operations “rollback” hides a consequential difference.
Decide who can authorise each operation. A technically possible return may still be unacceptable if it would discard work people have already completed. The owner needs that consequence in plain language, alongside the engineering options.
What should remain after the upgrade?
Before declaring the work complete, identify the pieces worth keeping. A compatibility check may belong in continuous integration. A recovery procedure needs a maintenance owner. Temporary dual-operation machinery needs an expiry condition, otherwise the transition becomes the permanent architecture.
The useful closing question is specific: could another engineer perform the next routine upgrade with less guesswork? If the answer depends on finding the person who remembers this one, there is still knowledge to transfer.
For your next migration review, add three columns to the plan: current operating state, evidence required to advance, and available recovery action. Fill in the first transition before debating the final date. Use the expedition brief to bound the work needed to make that transition credible.