Lesson 287 · AWS Learning Path

AWS 287: Migration waves, cutover and rollback

· Published · 12 min read

Labelled process diagram for AWS 287: Assessed applications and dependencies to Risk-balanced migration wave to Rehearsal and governed cutover to Validation, rollback or acceptance, and decommission, with decision...

Why this lesson matters

A migration wave is a managed production change, not a spreadsheet batch. It combines dependency groups only when the target platform, business calendar, people, replication capacity, validation and rollback paths can support them at the same time. A technically correct server migration can still fail because the DNS owner is absent, the service desk was not briefed, a business test has no approver, or target writes make rollback impossible.

This lesson joins the service-specific mechanisms in AWS283-AWS286 into a repeatable operating model. It emphasizes the evidence needed to say go, hold, fix forward, roll back, accept, and eventually decommission.

Outcomes

By the end, you can:

  • distinguish application, move/dependency group, wave, sprint and cutover event;
  • build risk-balanced waves without splitting hard dependencies;
  • calculate team and infrastructure concurrency instead of grouping by server count;
  • define readiness gates for business, platform, security, operations and data;
  • write an executable runbook with owner, command, expected result and rollback;
  • rehearse timing, communication, access and failure decisions;
  • run a command center with one UTC timeline and explicit authority;
  • choose all-at-once or phased traffic transition from consistency constraints;
  • handle DNS, sessions, queues, caches and post-cutover writes during rollback;
  • define hypercare, acceptance, source freeze and decommission evidence; and
  • produce a three-wave plan for 15 supplied applications.

Terms and hierarchy

TermMeaning
Applicationcomponents providing a business capability
Dependency/move groupapplications/components that must migrate together because of hard technical or nontechnical dependencies
Waveone or more move groups scheduled and governed as a cohort
Sprintplanning/build/test work period that prepares one or more waves
Cutover eventbounded window in which production authority moves to the target
Hypercareelevated post-cutover monitoring/support before normal operations accept ownership

One application can have several environments and strategies. Shared DNS, identity and monitoring are platform prerequisites, not reasons to place the entire portfolio in one move group.

Plan ahead, learn continuously

AWS large-migration guidance recommends planning several waves ahead so the migration factory has ready work, while treating wave planning as ongoing. Early waves should be small, lower-risk and useful for learning. Later waves can increase size and complexity only when measured results show the teams, platform and runbooks can handle it.

Avoid two extremes:

  • planning all 423 details months ahead as if dependencies and dates will not change; and
  • selecting next weekend's applications without enough lead time for remediation, testing and owner commitment.

Maintain a rolling horizon: committed near-term waves, prepared medium-term waves and candidate later waves. Revalidate new discovery, business calendars, application releases, platform changes and team availability at every commitment gate.

Build move groups before waves

Classify each dependency:

  • hard synchronous/latency dependency;
  • shared database or transaction boundary;
  • asynchronous queue/event/file exchange;
  • identity, DNS, PKI, proxy, licensing or platform dependency;
  • operational dependency such as deployment, backup or monitoring;
  • business, compliance, contract or shared-owner dependency;
  • bridgeable dependency with measured temporary connectivity; or
  • unknown dependency that blocks commitment.

Move together only when separation cannot meet measured requirements or transition risk. Document temporary hybrid links, who monitors them and when they are removed. An observed TCP connection alone does not create a move group; business meaning and tolerance do.

Wave selection criteria

Score transparently and retain raw evidence.

DimensionExample evidence
Businesscriticality, users, revenue, deadline, blackout, owner
Complexityservers, strategies, OS/database, state, target changes
Dependencyhard edges, external parties, shared service, latency
Datasize, change rate, consistency, seed/delta duration
Readinesstarget, remediation, tests, runbook, approval status
RecoveryRTO/RPO, rollback duration, write reconciliation
Peopleapp/platform/security/vendor/service-desk availability
Capacitybandwidth, quotas, replication jobs, test/cutover slots

A low total score must not hide a critical unknown. Set mandatory gates separately. A missing owner, unknown database link or untested rollback is a stop condition, not a few penalty points.

Size by constraints, not server count

Wave capacity is the minimum of several limits:

safe wave capacity = min(
  migration engineers,
  application validators,
  database/network/security change capacity,
  replication and network capacity,
  target quotas/capacity,
  command-center and service-desk support,
  rollback work that can fit before deadline
)

Ten stateless web servers may be easier than one 20-TB transactional database. Estimate effort per runbook step and team, then identify simultaneous tasks and bottleneck roles. Do not schedule two waves whose only DNS, DBA or approver is the same person.

Limit concurrent irreversible actions. Stagger move groups so the command center can detect and contain failure before the next traffic switch.

Wave phases and gates

Use explicit phases:

  1. Design: target, strategy, dependency and nonfunctional requirements approved.
  2. Build/pre-migration: account/network/security/identity/observability/backup and migration tooling ready.
  3. Test/rehearse: technical, business, performance, recovery and rollback evidence passes.
  4. Commit: owners, change, calendar, staffing, communications and entry criteria signed.
  5. Cutover: execute frozen runbook and decide go/hold/rollback/fix-forward.
  6. Hypercare: elevated observation, defect handling and cost/security review.
  7. Accept/decommission: operations accepts target; source is retained then retired through separate gates.

Each phase has entry evidence, exit evidence, accountable approver and a path back. A dashboard color is not approval.

The readiness dossier

Before commitment, require:

  • immutable scope: application/components/source IDs/target IDs;
  • current strategy, target architecture and dependency map;
  • target account/Region/VPC/DNS/security/identity/KMS readiness;
  • source/target backup and tested restore;
  • replication/full-load state, lag and expected final synchronization;
  • application, integration, data, performance, failover and security tests;
  • capacity, quotas, licenses, vendor support and cost owner;
  • approved runbook and rollback/fix-forward decision tree;
  • change ticket, freeze period and business blackout check;
  • named owners/alternates and access validated before the window;
  • user/service-desk/vendor communications; and
  • evidence repository, UTC clock and command-center details.

Use expiration dates. A test from six months and three releases ago may no longer prove readiness.

Write an executable runbook

Every step should include:

FieldPurpose
ID and planned UTC timeordering and timeline correlation
owner and alternateone accountable executor
preconditionevidence required before execution
exact actioncommand/console path/change reference with parameters reviewed
expected resultmeasurable evidence, not “looks good”
verification ownerseparation for critical steps
timeoutprevents indefinite waiting
on-failure actionretry, hold, fix forward or rollback step
evidence locationlog/screenshot/query/result without secrets
point-of-no-easy-return flagprompts formal approval

Avoid placeholders such as “team checks app.” State who runs which synthetic transaction, expected status/latency/data, and where evidence is recorded. Commands must be peer-reviewed and tested in a safe environment. Never paste secrets into the runbook.

Rehearsal and game day

Conduct a tabletop first, then a technical rehearsal using production-shaped infrastructure and representative data. Measure rather than estimate:

  • source freeze and final synchronization time;
  • target launch/deploy/restore duration;
  • DNS/load-balancer/route propagation and connection draining;
  • test duration and approver response;
  • rollback restoration and data reconciliation;
  • access to consoles, bastions, logs and vendor support;
  • communication/escalation latency; and
  • total time to irreversible decision deadline.

Inject failures: high replication lag, missing approver, certificate error, unhealthy targets, stale DNS, target performance regression and a post-cutover write. Update the runbook with observed timings and lessons. A rehearsal that cannot trigger rollback does not prove rollback.

Command-center roles

RoleAccountability
Cutover managerowns timeline, gates, holds and status cadence
Technical execution leadsrun platform/network/data/application steps
Evidence recorderrecords UTC actions/results and decision basis
Business/application owneraccepts functional service and business risk
Rollback authoritymakes deadline decision from predefined criteria
Security/compliancevalidates controls and handles findings
Operations/service deskmonitors incidents/users and accepts handoff
Communications leadsends approved stakeholder/customer updates

Use one bridge/chat/timeline and a separate escalation path. Executors report facts; the cutover manager controls sequence. Avoid side-channel changes that are absent from the log.

Pre-cutover freeze and baseline

Freeze the runbook, infrastructure/application versions and nonessential changes. Capture:

  • source health and key SLI baseline;
  • inventory/configuration hashes and deployment versions;
  • replication lag/checkpoint and last successful backup;
  • active users/sessions/jobs/queues and scheduled work;
  • DNS TTLs, current answers and load-balancer targets;
  • target health, alarms, synthetic tests and cost baseline; and
  • open incidents, accepted defects and rollback deadline.

Lower DNS TTL sufficiently before the event, but understand cache floors, negative caching and long-lived connections. TTL is not a global switch.

Cutover sequence

A typical stateful cutover is:

entry gate -> notify start -> stop schedulers/ingestion/writers
-> drain sessions/queues -> final backup/checkpoint
-> final sync and prove lag zero/threshold
-> activate target -> switch routing/DNS/integrations
-> technical smoke tests -> business transactions
-> observe SLIs and data -> accept, fix forward, hold, or roll back

The actual order depends on architecture. For some systems target services start before source freeze; for others duplicate identity/licenses require all-at-once activation. Document the source of truth and write authority at every minute.

All-at-once versus phased traffic

All-at-once can be safer where simultaneous active systems cause hostname, license, session or data conflicts. It concentrates risk and demands strong rehearsal.

Phased/canary transition reduces blast radius and provides early evidence, but requires traffic partitioning, compatible data access and comparable user cohorts. Weighted DNS alone is not deterministic per request and cached resolvers can distort percentages. Load balancer weights, headers, tenants or feature flags may provide better control.

Do not phase writers across databases unless consistency and conflict handling are proven.

Objective acceptance and rollback triggers

Examples:

  • HTTP successful transaction rate below 99.9 percent for five minutes;
  • p95 latency above 800 ms for ten minutes versus 350 ms baseline;
  • any financial reconciliation mismatch;
  • replication not through checkpoint by T-minus 60;
  • more than two critical integrations failing;
  • restore/backup/security controls not active;
  • cutover behind schedule such that rollback cannot finish before outage end; or
  • business owner unavailable at final gate.

Set severity, observation window, data source, owner and action. Avoid “rollback if serious issue.” Preserve fix-forward budget: a small known configuration defect may be faster/safer to repair, while data corruption should trigger an immediate stop.

Rollback before target writes

If target has received no authoritative writes, rollback can often:

  1. stop target traffic/workers;
  2. restore routes/DNS/load-balancer to source;
  3. restart source ingestion/jobs in dependency order;
  4. verify source health and transactions;
  5. communicate restoration; and
  6. preserve target/failure evidence for analysis.

Even then account for DNS caches, sessions, messages and duplicate automation. Prove that source remained intact and current.

Rollback after target writes

Once target accepts writes, old source is stale. “Change DNS back” can lose orders or duplicate messages. Choose and rehearse one:

  • keep target read-only until acceptance;
  • reverse replication/fail-forward database prepared before cutover;
  • dual write with idempotency, conflict resolution and observability;
  • transaction journal/export/replay back to source;
  • native backup/restore within measured RTO; or
  • remain on target and fix forward because rollback data risk is greater.

Define last safe rollback time and point of no return. Reconcile transaction IDs, amounts/counts, queue offsets, files and side effects such as email/payment before reopening. A rollback can take longer than cutover.

DNS, sessions, caches and integrations

Traffic transition includes more than one DNS record:

  • authoritative records and TTL/negative cache;
  • resolver and client caches;
  • persistent TCP/TLS sessions and connection pools;
  • load-balancer deregistration/draining;
  • API allowlists, partner endpoints and fixed IPs;
  • certificates/SNI, redirects and callback URLs;
  • queue consumers, cron/batch and webhooks;
  • CDN/cache invalidation and stale content; and
  • health checks that may not represent business function.

Test from multiple networks and resolver paths. Preserve old routes/records in the runbook for rollback without leaving an uncontrolled bypass.

Hypercare

Hypercare is a defined state, not “watch it closely.” Specify duration and exit criteria:

  • SLO/SLI within baseline for representative peaks;
  • no unresolved critical/high migration defects;
  • data reconciliation and scheduled cycles pass;
  • backups complete and restore evidence exists;
  • security/operations findings assigned and controlled;
  • cost and capacity match expected ranges;
  • alerts, dashboards, on-call and runbooks accepted by operations;
  • users/service desk see no unexplained pattern; and
  • known issues, technical debt and optimization backlog transferred.

Use follow-the-sun coverage where required. Hold daily defect/risk reviews and retain migration team support until formal handoff.

Acceptance is not decommissioning

After application acceptance, freeze old source from ordinary use but retain it according to rollback, audit and data-retention policy. Then separately approve:

  • archive/backup and tested retrieval;
  • no remaining callers, DNS, jobs or integrations;
  • legal hold and data destruction requirements;
  • licenses/contracts and monitoring closure;
  • CMDB/asset/security inventory update;
  • credentials/certificates/firewall/routes cleanup;
  • replication/tool/test resource cleanup; and
  • final billing confirmation.

Make decommissioning reversible during an observation period where possible. Record what cannot be recovered after destruction.

Migration factory metrics and learning

Track more than servers per week:

  • waves committed/completed/rolled back/deferred;
  • readiness defects found before versus during cutover;
  • planned versus actual duration by step;
  • migration-caused incidents and business impact;
  • validation first-pass rate and unresolved exceptions;
  • replication/capacity bottlenecks;
  • rollback rehearsal and actual recovery time;
  • hypercare defects, duration and operations acceptance;
  • cost per application/move group; and
  • repeated runbook defects automated or eliminated.

Velocity that increases incident or rollback risk is not success. Feed lessons into future-wave criteria, automation and training without changing an already frozen event unnoticed.

Troubleshooting the program

SymptomLikely causeResponse
Waves repeatedly deferreadiness criteria applied too late or portfolio data stalemove gates earlier; assign gap owners and expiry
Cutovers exceed windowsunmeasured sync/tests, too much concurrency, bottleneck roleuse actual timing; resize/stagger wave
Green dashboards but user failureinfrastructure-only testsadd business transactions and external-path probes
Frequent DNS rollback issuesTTL/cache/session behavior assumedtest resolver/client/draining paths and maintain explicit records
Rollback impossible after writesno write-authority/reconciliation designstop cutover; implement/test data strategy first
Hypercare never endsno exit criteria or operations handoffdefine SLO/defect/backup/security/cost gates and owners
Source costs remainacceptance confused with decommissionexecute separate retained-source and cleanup governance

Hands-on workshop: 15 applications, three waves

Produce:

  1. application/dependency inventory with hard, soft and unknown edges;
  2. move groups and reasons;
  3. business/complexity/readiness/recovery evidence;
  4. team/capacity constraint calculation;
  5. three waves with rejection/defer reasons;
  6. phase gates and readiness dossier;
  7. a 60-step minute-by-minute runbook for one stateful move group;
  8. RACI, contact/access check and communications matrix;
  9. rehearsal report with measured times and injected failures;
  10. SLI/data/DNS acceptance and rollback thresholds;
  11. target-write rollback/fix-forward design;
  12. hypercare and operations handoff criteria; and
  13. retained-source/decommission checklist and cost forecast.

Inject: unknown payment dependency, DBA double-booking, lag above threshold, stale DNS at one branch, certificate mismatch, target p95 regression, and 23 target-only orders. The plan must show exactly where execution stops and who decides.

Knowledge check

  1. Why is a move group not a wave?

A move group expresses co-dependency; a wave combines move groups according to schedule, capacity and risk.

  1. Why can a low risk score still block migration?

Mandatory unknowns such as no owner or untested rollback cannot be averaged away.

  1. What is the key rollback boundary?

Whether the target has accepted authoritative writes that the source does not contain.

  1. Why is lowering DNS TTL insufficient?

Resolver/client caches, negative caching, persistent sessions and integration allowlists still affect traffic.

  1. What makes a runbook executable?

Named owner, exact action, precondition, measurable result, timeout, failure action and retained evidence.

  1. When does hypercare end?

When predefined service, data, backup, security, cost, defect and operations-handoff criteria pass.

  1. Does application acceptance permit immediate source deletion?

No. Decommissioning has separate rollback, retention, legal, dependency and owner gates.

Lesson acceptance

Pass only if every move group/wave decision, capacity assumption, runbook action, gate, metric, authority, target-write treatment, hypercare exit and decommission step is reproducible. Reject plans grouped by server count alone, dependent on one unavailable specialist, using vague tests/triggers, treating DNS as instantaneous, rolling back stale data, deleting source at acceptance, or reporting velocity without risk/outcome measures.

Official sources

Advertisement