AWS 287: Migration waves, cutover and rollback
Why this lesson matters
A migration wave is a managed production change, not a spreadsheet batch. It combines dependency groups only when the target platform, business calendar, people, replication capacity, validation and rollback paths can support them at the same time. A technically correct server migration can still fail because the DNS owner is absent, the service desk was not briefed, a business test has no approver, or target writes make rollback impossible.
This lesson joins the service-specific mechanisms in AWS283-AWS286 into a repeatable operating model. It emphasizes the evidence needed to say go, hold, fix forward, roll back, accept, and eventually decommission.
Outcomes
By the end, you can:
- distinguish application, move/dependency group, wave, sprint and cutover event;
- build risk-balanced waves without splitting hard dependencies;
- calculate team and infrastructure concurrency instead of grouping by server count;
- define readiness gates for business, platform, security, operations and data;
- write an executable runbook with owner, command, expected result and rollback;
- rehearse timing, communication, access and failure decisions;
- run a command center with one UTC timeline and explicit authority;
- choose all-at-once or phased traffic transition from consistency constraints;
- handle DNS, sessions, queues, caches and post-cutover writes during rollback;
- define hypercare, acceptance, source freeze and decommission evidence; and
- produce a three-wave plan for 15 supplied applications.
Terms and hierarchy
| Term | Meaning |
|---|---|
| Application | components providing a business capability |
| Dependency/move group | applications/components that must migrate together because of hard technical or nontechnical dependencies |
| Wave | one or more move groups scheduled and governed as a cohort |
| Sprint | planning/build/test work period that prepares one or more waves |
| Cutover event | bounded window in which production authority moves to the target |
| Hypercare | elevated post-cutover monitoring/support before normal operations accept ownership |
One application can have several environments and strategies. Shared DNS, identity and monitoring are platform prerequisites, not reasons to place the entire portfolio in one move group.
Plan ahead, learn continuously
AWS large-migration guidance recommends planning several waves ahead so the migration factory has ready work, while treating wave planning as ongoing. Early waves should be small, lower-risk and useful for learning. Later waves can increase size and complexity only when measured results show the teams, platform and runbooks can handle it.
Avoid two extremes:
- planning all 423 details months ahead as if dependencies and dates will not change; and
- selecting next weekend's applications without enough lead time for remediation, testing and owner commitment.
Maintain a rolling horizon: committed near-term waves, prepared medium-term waves and candidate later waves. Revalidate new discovery, business calendars, application releases, platform changes and team availability at every commitment gate.
Build move groups before waves
Classify each dependency:
- hard synchronous/latency dependency;
- shared database or transaction boundary;
- asynchronous queue/event/file exchange;
- identity, DNS, PKI, proxy, licensing or platform dependency;
- operational dependency such as deployment, backup or monitoring;
- business, compliance, contract or shared-owner dependency;
- bridgeable dependency with measured temporary connectivity; or
- unknown dependency that blocks commitment.
Move together only when separation cannot meet measured requirements or transition risk. Document temporary hybrid links, who monitors them and when they are removed. An observed TCP connection alone does not create a move group; business meaning and tolerance do.
Wave selection criteria
Score transparently and retain raw evidence.
| Dimension | Example evidence |
|---|---|
| Business | criticality, users, revenue, deadline, blackout, owner |
| Complexity | servers, strategies, OS/database, state, target changes |
| Dependency | hard edges, external parties, shared service, latency |
| Data | size, change rate, consistency, seed/delta duration |
| Readiness | target, remediation, tests, runbook, approval status |
| Recovery | RTO/RPO, rollback duration, write reconciliation |
| People | app/platform/security/vendor/service-desk availability |
| Capacity | bandwidth, quotas, replication jobs, test/cutover slots |
A low total score must not hide a critical unknown. Set mandatory gates separately. A missing owner, unknown database link or untested rollback is a stop condition, not a few penalty points.
Size by constraints, not server count
Wave capacity is the minimum of several limits:
safe wave capacity = min(
migration engineers,
application validators,
database/network/security change capacity,
replication and network capacity,
target quotas/capacity,
command-center and service-desk support,
rollback work that can fit before deadline
)
Ten stateless web servers may be easier than one 20-TB transactional database. Estimate effort per runbook step and team, then identify simultaneous tasks and bottleneck roles. Do not schedule two waves whose only DNS, DBA or approver is the same person.
Limit concurrent irreversible actions. Stagger move groups so the command center can detect and contain failure before the next traffic switch.
Wave phases and gates
Use explicit phases:
- Design: target, strategy, dependency and nonfunctional requirements approved.
- Build/pre-migration: account/network/security/identity/observability/backup and migration tooling ready.
- Test/rehearse: technical, business, performance, recovery and rollback evidence passes.
- Commit: owners, change, calendar, staffing, communications and entry criteria signed.
- Cutover: execute frozen runbook and decide go/hold/rollback/fix-forward.
- Hypercare: elevated observation, defect handling and cost/security review.
- Accept/decommission: operations accepts target; source is retained then retired through separate gates.
Each phase has entry evidence, exit evidence, accountable approver and a path back. A dashboard color is not approval.
The readiness dossier
Before commitment, require:
- immutable scope: application/components/source IDs/target IDs;
- current strategy, target architecture and dependency map;
- target account/Region/VPC/DNS/security/identity/KMS readiness;
- source/target backup and tested restore;
- replication/full-load state, lag and expected final synchronization;
- application, integration, data, performance, failover and security tests;
- capacity, quotas, licenses, vendor support and cost owner;
- approved runbook and rollback/fix-forward decision tree;
- change ticket, freeze period and business blackout check;
- named owners/alternates and access validated before the window;
- user/service-desk/vendor communications; and
- evidence repository, UTC clock and command-center details.
Use expiration dates. A test from six months and three releases ago may no longer prove readiness.
Write an executable runbook
Every step should include:
| Field | Purpose |
|---|---|
| ID and planned UTC time | ordering and timeline correlation |
| owner and alternate | one accountable executor |
| precondition | evidence required before execution |
| exact action | command/console path/change reference with parameters reviewed |
| expected result | measurable evidence, not “looks good” |
| verification owner | separation for critical steps |
| timeout | prevents indefinite waiting |
| on-failure action | retry, hold, fix forward or rollback step |
| evidence location | log/screenshot/query/result without secrets |
| point-of-no-easy-return flag | prompts formal approval |
Avoid placeholders such as “team checks app.” State who runs which synthetic transaction, expected status/latency/data, and where evidence is recorded. Commands must be peer-reviewed and tested in a safe environment. Never paste secrets into the runbook.
Rehearsal and game day
Conduct a tabletop first, then a technical rehearsal using production-shaped infrastructure and representative data. Measure rather than estimate:
- source freeze and final synchronization time;
- target launch/deploy/restore duration;
- DNS/load-balancer/route propagation and connection draining;
- test duration and approver response;
- rollback restoration and data reconciliation;
- access to consoles, bastions, logs and vendor support;
- communication/escalation latency; and
- total time to irreversible decision deadline.
Inject failures: high replication lag, missing approver, certificate error, unhealthy targets, stale DNS, target performance regression and a post-cutover write. Update the runbook with observed timings and lessons. A rehearsal that cannot trigger rollback does not prove rollback.
Command-center roles
| Role | Accountability |
|---|---|
| Cutover manager | owns timeline, gates, holds and status cadence |
| Technical execution leads | run platform/network/data/application steps |
| Evidence recorder | records UTC actions/results and decision basis |
| Business/application owner | accepts functional service and business risk |
| Rollback authority | makes deadline decision from predefined criteria |
| Security/compliance | validates controls and handles findings |
| Operations/service desk | monitors incidents/users and accepts handoff |
| Communications lead | sends approved stakeholder/customer updates |
Use one bridge/chat/timeline and a separate escalation path. Executors report facts; the cutover manager controls sequence. Avoid side-channel changes that are absent from the log.
Pre-cutover freeze and baseline
Freeze the runbook, infrastructure/application versions and nonessential changes. Capture:
- source health and key SLI baseline;
- inventory/configuration hashes and deployment versions;
- replication lag/checkpoint and last successful backup;
- active users/sessions/jobs/queues and scheduled work;
- DNS TTLs, current answers and load-balancer targets;
- target health, alarms, synthetic tests and cost baseline; and
- open incidents, accepted defects and rollback deadline.
Lower DNS TTL sufficiently before the event, but understand cache floors, negative caching and long-lived connections. TTL is not a global switch.
Cutover sequence
A typical stateful cutover is:
entry gate -> notify start -> stop schedulers/ingestion/writers
-> drain sessions/queues -> final backup/checkpoint
-> final sync and prove lag zero/threshold
-> activate target -> switch routing/DNS/integrations
-> technical smoke tests -> business transactions
-> observe SLIs and data -> accept, fix forward, hold, or roll back
The actual order depends on architecture. For some systems target services start before source freeze; for others duplicate identity/licenses require all-at-once activation. Document the source of truth and write authority at every minute.
All-at-once versus phased traffic
All-at-once can be safer where simultaneous active systems cause hostname, license, session or data conflicts. It concentrates risk and demands strong rehearsal.
Phased/canary transition reduces blast radius and provides early evidence, but requires traffic partitioning, compatible data access and comparable user cohorts. Weighted DNS alone is not deterministic per request and cached resolvers can distort percentages. Load balancer weights, headers, tenants or feature flags may provide better control.
Do not phase writers across databases unless consistency and conflict handling are proven.
Objective acceptance and rollback triggers
Examples:
- HTTP successful transaction rate below 99.9 percent for five minutes;
- p95 latency above 800 ms for ten minutes versus 350 ms baseline;
- any financial reconciliation mismatch;
- replication not through checkpoint by T-minus 60;
- more than two critical integrations failing;
- restore/backup/security controls not active;
- cutover behind schedule such that rollback cannot finish before outage end; or
- business owner unavailable at final gate.
Set severity, observation window, data source, owner and action. Avoid “rollback if serious issue.” Preserve fix-forward budget: a small known configuration defect may be faster/safer to repair, while data corruption should trigger an immediate stop.
Rollback before target writes
If target has received no authoritative writes, rollback can often:
- stop target traffic/workers;
- restore routes/DNS/load-balancer to source;
- restart source ingestion/jobs in dependency order;
- verify source health and transactions;
- communicate restoration; and
- preserve target/failure evidence for analysis.
Even then account for DNS caches, sessions, messages and duplicate automation. Prove that source remained intact and current.
Rollback after target writes
Once target accepts writes, old source is stale. “Change DNS back” can lose orders or duplicate messages. Choose and rehearse one:
- keep target read-only until acceptance;
- reverse replication/fail-forward database prepared before cutover;
- dual write with idempotency, conflict resolution and observability;
- transaction journal/export/replay back to source;
- native backup/restore within measured RTO; or
- remain on target and fix forward because rollback data risk is greater.
Define last safe rollback time and point of no return. Reconcile transaction IDs, amounts/counts, queue offsets, files and side effects such as email/payment before reopening. A rollback can take longer than cutover.
DNS, sessions, caches and integrations
Traffic transition includes more than one DNS record:
- authoritative records and TTL/negative cache;
- resolver and client caches;
- persistent TCP/TLS sessions and connection pools;
- load-balancer deregistration/draining;
- API allowlists, partner endpoints and fixed IPs;
- certificates/SNI, redirects and callback URLs;
- queue consumers, cron/batch and webhooks;
- CDN/cache invalidation and stale content; and
- health checks that may not represent business function.
Test from multiple networks and resolver paths. Preserve old routes/records in the runbook for rollback without leaving an uncontrolled bypass.
Hypercare
Hypercare is a defined state, not “watch it closely.” Specify duration and exit criteria:
- SLO/SLI within baseline for representative peaks;
- no unresolved critical/high migration defects;
- data reconciliation and scheduled cycles pass;
- backups complete and restore evidence exists;
- security/operations findings assigned and controlled;
- cost and capacity match expected ranges;
- alerts, dashboards, on-call and runbooks accepted by operations;
- users/service desk see no unexplained pattern; and
- known issues, technical debt and optimization backlog transferred.
Use follow-the-sun coverage where required. Hold daily defect/risk reviews and retain migration team support until formal handoff.
Acceptance is not decommissioning
After application acceptance, freeze old source from ordinary use but retain it according to rollback, audit and data-retention policy. Then separately approve:
- archive/backup and tested retrieval;
- no remaining callers, DNS, jobs or integrations;
- legal hold and data destruction requirements;
- licenses/contracts and monitoring closure;
- CMDB/asset/security inventory update;
- credentials/certificates/firewall/routes cleanup;
- replication/tool/test resource cleanup; and
- final billing confirmation.
Make decommissioning reversible during an observation period where possible. Record what cannot be recovered after destruction.
Migration factory metrics and learning
Track more than servers per week:
- waves committed/completed/rolled back/deferred;
- readiness defects found before versus during cutover;
- planned versus actual duration by step;
- migration-caused incidents and business impact;
- validation first-pass rate and unresolved exceptions;
- replication/capacity bottlenecks;
- rollback rehearsal and actual recovery time;
- hypercare defects, duration and operations acceptance;
- cost per application/move group; and
- repeated runbook defects automated or eliminated.
Velocity that increases incident or rollback risk is not success. Feed lessons into future-wave criteria, automation and training without changing an already frozen event unnoticed.
Troubleshooting the program
| Symptom | Likely cause | Response |
|---|---|---|
| Waves repeatedly defer | readiness criteria applied too late or portfolio data stale | move gates earlier; assign gap owners and expiry |
| Cutovers exceed windows | unmeasured sync/tests, too much concurrency, bottleneck role | use actual timing; resize/stagger wave |
| Green dashboards but user failure | infrastructure-only tests | add business transactions and external-path probes |
| Frequent DNS rollback issues | TTL/cache/session behavior assumed | test resolver/client/draining paths and maintain explicit records |
| Rollback impossible after writes | no write-authority/reconciliation design | stop cutover; implement/test data strategy first |
| Hypercare never ends | no exit criteria or operations handoff | define SLO/defect/backup/security/cost gates and owners |
| Source costs remain | acceptance confused with decommission | execute separate retained-source and cleanup governance |
Hands-on workshop: 15 applications, three waves
Produce:
- application/dependency inventory with hard, soft and unknown edges;
- move groups and reasons;
- business/complexity/readiness/recovery evidence;
- team/capacity constraint calculation;
- three waves with rejection/defer reasons;
- phase gates and readiness dossier;
- a 60-step minute-by-minute runbook for one stateful move group;
- RACI, contact/access check and communications matrix;
- rehearsal report with measured times and injected failures;
- SLI/data/DNS acceptance and rollback thresholds;
- target-write rollback/fix-forward design;
- hypercare and operations handoff criteria; and
- retained-source/decommission checklist and cost forecast.
Inject: unknown payment dependency, DBA double-booking, lag above threshold, stale DNS at one branch, certificate mismatch, target p95 regression, and 23 target-only orders. The plan must show exactly where execution stops and who decides.
Knowledge check
- Why is a move group not a wave?
A move group expresses co-dependency; a wave combines move groups according to schedule, capacity and risk.
- Why can a low risk score still block migration?
Mandatory unknowns such as no owner or untested rollback cannot be averaged away.
- What is the key rollback boundary?
Whether the target has accepted authoritative writes that the source does not contain.
- Why is lowering DNS TTL insufficient?
Resolver/client caches, negative caching, persistent sessions and integration allowlists still affect traffic.
- What makes a runbook executable?
Named owner, exact action, precondition, measurable result, timeout, failure action and retained evidence.
- When does hypercare end?
When predefined service, data, backup, security, cost, defect and operations-handoff criteria pass.
- Does application acceptance permit immediate source deletion?
No. Decommissioning has separate rollback, retention, legal, dependency and owner gates.
Lesson acceptance
Pass only if every move group/wave decision, capacity assumption, runbook action, gate, metric, authority, target-write treatment, hypercare exit and decommission step is reproducible. Reject plans grouped by server count alone, dependent on one unavailable specialist, using vague tests/triggers, treating DNS as instantaneous, rolling back stale data, deleting source at acceptance, or reporting velocity without risk/outcome measures.