AWS 292: Architecture: build a migration wave, cutover, rollback, and validation plan
Why this lesson matters
A migration can be technically possible and still fail during cutover because no one can answer four questions quickly: What happens next? What proves it worked? Who can decide? How do we return safely?
This capstone turns the outputs of P16 into an executable decision package for one business service. It is not another service overview. The learner must connect discovery, strategy, business case, target design, landing-zone readiness, replication, testing, change control, traffic movement, data ownership, rollback, business acceptance, hypercare, and decommissioning.
The result must be usable by people who did not write it, at an inconvenient hour, while facts change. Ambiguous verbs such as “validate,” “monitor,” and “rollback if required” are replaced with commands, expected results, deadlines, owners, and decision authority.
What you will be able to do
By the end, you can:
- distinguish an architecture decision record, wave plan, cutover runbook, and evidence log;
- prove that a dependency group is safe to move in one window;
- build readiness gates that cannot be passed by opinion alone;
- calculate the decision deadline from downtime and rollback duration;
- define pre-write and post-write rollback branches;
- sequence technical, business, communication, and change-management tasks;
- validate infrastructure, security, data, integration, performance, recovery, and business outcomes;
- run a command center with explicit authority and escalation;
- transfer ownership through hypercare without silently decommissioning the source; and
- defend the complete package in an architecture review.
Before you start
- Use the supplied fictional service. Do not run a migration against a real workload.
- Existing-account inspection requires owner approval and remains read-only. A source or target identifier can expose sensitive architecture; redact it from shared evidence.
- Never stop a server, change DNS, disable a job, initiate DMS or MGN actions, execute a DataSync task, or alter routes while completing this lesson.
- Do not place passwords, keys, tokens, private hostnames, customer records, or full account IDs in the runbook. Reference an approved secret and access process.
- All times in the final runbook must use one explicit time zone and include the UTC equivalent.
- A plan is not approved merely because a tool says replication is healthy. Named technical, business, security, operations, and change owners must accept their gates.
1. Know the four artifacts
| Artifact | Question answered | Change frequency | Approval |
|---|---|---|---|
| Architecture decision record (ADR) | Why this scope, target, strategy, and risk treatment were chosen | When material assumptions or design change | Architecture and accountable business/technology owners |
| Wave plan | Which dependency groups move together, through which phases, and when | Rolling planning cycle | Program and affected teams |
| Cutover runbook | Exactly who does what, when, using which evidence and recovery step | Rehearsal findings and controlled updates | Cutover/change authority |
| Execution and evidence log | What actually occurred, result, timestamp, approver, and deviation | During and after execution | Immutable or controlled operational record |
Do not hide operational detail inside the ADR. Do not turn a live runbook into a design essay. Link artifacts by stable version and change record so the command center knows which approved version is authoritative.
2. The supplied business service
Build the capstone for Orchid Orders, a fictional three-tier order service:
- two Linux web servers behind an on-premises load balancer;
- two Java application servers with in-memory sessions;
- PostgreSQL primary plus reporting replica;
- a file exchange used by warehouse and finance systems;
- outbound payment, identity, email, and carrier APIs;
- nightly settlement at 01:00 local time; and
- 24x7 customer ordering with the lowest demand Saturday 22:00-00:00.
The approved direction is rehost web and application servers using AWS Application Migration Service, replatform PostgreSQL to Amazon RDS for PostgreSQL using AWS DMS full load plus change data capture, and move file exchange to an approved AWS storage target using DataSync. Traffic enters through the approved target load-balancing and DNS design. This is a teaching scenario, not a claim that these services fit every workload.
Requirements:
- maximum approved customer outage: 90 minutes;
- cutover window: 120 minutes, including the decision and rollback reserve;
- RTO: 120 minutes; RPO at committed go-live: zero acknowledged orders lost;
- target runs across two Availability Zones where the selected service supports it;
- database and application tiers are private;
- payment and warehouse integrations must pass before go-live;
- a business owner must validate one test order through settlement-visible state;
- source infrastructure must remain recoverable for 30 days; and
- production source disposal requires separate retention, finance, security, and application-owner approval.
Known complications include fixed source IP allowlists, a 15-minute DNS TTL, in-memory sessions, one non-idempotent batch submission, a reporting user with an old PostgreSQL extension, and a carrier partner available only until 23:15.
3. Fix scope and the consistency boundary
A wave is a management grouping. A move group is the smallest set that must change together because splitting it would break business behavior or data consistency. One wave can contain multiple move groups with separate cutover decisions.
For every component record:
- current and target identity;
- application and infrastructure owner;
- strategy and migration tool;
- inbound and outbound dependencies;
- data read/write behavior and system of record;
- startup, shutdown, and health sequence;
- maintenance window and blackout dates;
- expected replication or transfer duration;
- rollback method and maximum duration; and
- source disposition.
Draw a dependency graph with direction, protocol, port, DNS name, identity, data classification, timeout, retry behavior, and owner. A CMDB relationship is a hypothesis until a representative observation window and an owner confirm it. Include scheduled, month-end, reporting, administrator, monitoring, backup, certificate, directory, and vendor flows that runtime discovery might miss.
Define the transactional consistency boundary. For Orchid Orders, order records, payment state, inventory request, and settlement event cannot be validated as unrelated servers. The plan must state where writes stop, how queues drain, how DMS latency reaches the accepted threshold, which database becomes authoritative, and how an accepted post-cutover order returns during rollback.
4. Write the ADR before the runbook
Use this ADR structure:
- Title, status, date, owners, reviewers, and superseded decisions.
- Decision statement in one paragraph.
- Business outcomes and measurable technical requirements.
- Scope, exclusions, move groups, and consistency boundary.
- Current architecture and evidence quality.
- Target architecture, accounts, Regions, Availability Zones, and trust boundaries.
- At least two viable alternatives.
- Selected strategy and migration pattern per component.
- Decision drivers, constraints, assumptions, and unknowns.
- Security, operations, resilience, data, and compliance consequences.
- Cost and schedule consequences, including parallel run.
- Risks, controls, owners, deadlines, and residual acceptance.
- Reversibility and point-of-no-simple-return.
- Validation, hypercare, decommission, and benefit-measurement approach.
- Approval, expiry, and revisit triggers.
Compare at least these options:
- move the complete service in one window;
- move read-only/reporting capability first and transactional capability later; and
- retain the service temporarily while closing the IP, session, extension, and idempotency gaps.
Use weighted scoring only as decision support. A high score cannot override a mandatory security, licensing, RTO, or data-correctness requirement. Document why rejected options were viable and why they lost.
5. Convert readiness into evidence gates
Every gate needs an ID, criterion, evidence location, evidence time, owner, approver, status, expiry, and waiver authority. “Green” is not evidence.
Foundation gate
Prove target accounts and ownership, VPC/subnet capacity, routes, DNS, time synchronization, logging, security services, KMS key access, backup, quotas, monitoring destinations, support model, break-glass access, tagging, budgets, and configuration baselines. Verify both normal operator and emergency access before the window.
Workload gate
Prove target builds, immutable versions, patches, certificates, secrets, service identities, least privilege, security groups, dependency connectivity, storage, startup order, health checks, scaling, backup, restore, and operational runbooks. A successful ping does not prove TLS name, application authorization, transaction, or return path.
Data gate
Record initial-load completion, CDC state, source-log retention, replication latency trend, error tables, schema exceptions, LOB treatment, row or object reconciliation, checksums where meaningful, time zones, sequences, triggers, and target write protection. Document who can declare the final synchronization complete.
Test gate
Require functional, integration, security, performance, failover/recovery, monitoring, backup/restore, batch, operational, and business acceptance evidence. Compare target performance with a dated source baseline under comparable load. Record defects and accepted residual risk rather than deleting failed results.
People and change gate
Prove approved change, maintenance communication, participant availability, vendor contacts, access tests, bridge details, escalation tree, support tickets, freeze scope, duty limits, backup personnel, and go/no-go authority. Confirm that every named person accepted the role.
Waivers must identify the unmet criterion, reason, compensating control, risk owner, expiry, and approver. A waiver does not turn failed evidence into passed evidence.
6. Design backward from the deadline
The rollback decision must occur early enough to finish rollback before the approved outage ends.
latest rollback decision = outage end
- proven rollback duration
- rollback validation duration
- contingency reserve
For a 90-minute outage, if tested rollback takes 25 minutes, source validation takes 10 minutes, and reserve is 10 minutes, the latest decision is minute 45. A plan that starts business testing at minute 50 has already surrendered a safe rollback.
Build a countdown plan from at least T-30 days, T-14, T-7, T-3, T-1, T-4 hours, and T-30 minutes. Include TTL reduction early enough for caches to age, change approvals, final rehearsal, freeze, replication health, participant confirmation, support readiness, backups, and abort conditions. Restore normal TTL only after stability and rollback strategy allow it.
7. Build an executable runbook
Each task row must include:
| Field | Required content |
|---|---|
| ID and phase | Stable reference and preparation, cutover, rollback, or recovery phase |
| Planned start/end | Absolute time plus expected duration |
| Predecessor | Task or gate that must succeed first |
| Executor and verifier | Two roles where separation matters |
| Exact action | Approved command, console path, automation job, or communication text |
| Expected result | Observable value, state, count, or threshold |
| Evidence | Timestamped output or controlled link |
| Failure action | Retry once, investigate within limit, skip by waiver, stop, or roll back |
| Rollback counterpart | Task restoring or compensating this change |
| Actuals | Start, finish, result, deviation, and incident/change reference |
Use explicit commands only after they have been tested in the same pattern and redacted for distribution. Avoid destructive wildcard operations and mutable “latest” artifacts. Every automation step needs inputs, version, dry-run or rehearsal evidence, idempotency behavior, timeout, output, and a manual recovery path.
Example timeline
21:30 bridge open, attendance and authority check
21:35 freeze verification and source health baseline
21:40 stop batch/schedulers; verify no new launches
21:45 place application in controlled write drain
21:50 drain sessions/queues; capture outstanding work
21:55 final file delta and database CDC catch-up
22:05 source writes stopped; consistency checkpoint recorded
22:10 final validation; GO-1 authorizes target write enablement
22:15 launch/enable target and perform technical smoke tests
22:25 shift controlled traffic and run integration tests
22:35 business order test and data reconciliation
22:45 final GO or rollback decision deadline
22:50 increase traffic only after GO
23:00 announce service restoration and begin hypercare
This example is incomplete until the learner supplies commands, owners, thresholds, rollback counterparts, and exact evidence. Parallel tasks must be visually identified; never assume row order alone conveys concurrency.
8. Establish command-center authority
Name these roles even if one person fills several:
- cutover commander, who controls sequence and bridge discipline;
- change authority, who confirms the approved window and material deviations;
- technical migration lead;
- application owner;
- business decision owner;
- source infrastructure, target cloud, network/DNS, database/data, security, and operations leads;
- test coordinator and evidence recorder;
- communications lead; and
- incident manager if service impact exceeds the plan.
Only one role issues GO, HOLD, FIX FORWARD, and ROLLBACK calls. Technical teams recommend; the named accountable decision owner decides within the deadline. Predetermine what happens if that person is unreachable.
Use a concise cadence: task ID, expected result, actual result, variance, risk to deadline, recommendation, decision. Side conversations must return decisions to the main log. Never allow optimistic verbal status to replace evidence.
9. Define objective decision states
Use four states:
- GO: all mandatory gates pass and remaining risk is accepted.
- HOLD: no irreversible step; investigate within a fixed budget.
- FIX FORWARD: target has writes or reverting is riskier, and repair can meet a defined deadline.
- ROLLBACK: a trigger fired and enough time remains for the tested path.
Example rollback triggers:
- mandatory payment or warehouse integration fails twice after one approved repair;
- reconciliation exceeds the approved count or value tolerance;
- sustained error rate, latency, or saturation exceeds a named threshold;
- target database cannot accept writes safely;
- security control or audit logging is absent;
- business test cannot complete before the decision deadline; or
- remaining rollback time falls below the proven requirement.
Avoid “major issue” or “performance unacceptable.” Those phrases transfer the decision to a debate under pressure.
10. Separate rollback before and after target writes
Before target writes
If the source remains authoritative and unchanged, rollback can stop target traffic, disable target processing, restore source schedulers and routes, return traffic, validate source operation, and announce restoration. Still account for DNS caches, sessions, queued messages, and external allowlists.
After target writes
Once orders are accepted in the target, the source is stale. Choose and rehearse one approach:
- replicate target changes back to a prepared fail-forward source;
- restore and replay an authoritative transaction or event log;
- use a deliberately engineered dual-write pattern with reconciliation;
- compensate business transactions under approved rules; or
- remain on the target and fix forward.
Specify conflict handling, order preservation, sequence values, duplicate prevention, encryption, transfer time, validation, and the owner who accepts any loss boundary. Native backup and restore is useful only when tested restore time and data point satisfy the deadline.
Define the point-of-no-simple-return. This can be first target write, irreversible schema change, partner allowlist removal, destructive source change, or replication finalization. Require a specific authorization immediately before it.
11. Build a validation matrix
Validation must test a chain, not isolated green dashboards.
| Layer | Example evidence | Acceptance |
|---|---|---|
| Infrastructure | Target state, AZ placement, capacity, quota, volume and load balancer health | Matches approved design and headroom |
| Network/DNS | Resolution from each client class, TLS, forward and return paths, partner source IP | Every required flow succeeds; forbidden flow fails |
| Identity/security | Workload identity, least privilege, secret retrieval, encryption, audit events | Allowed and denied cases behave as designed |
| Data | Counts, control totals, sampled records, sequence/encoding/time-zone checks, DMS exceptions | Within signed tolerance with no unexplained error |
| Application | Login, create/read/update, session, file, API, and error behavior | Critical journeys pass |
| Integration | Payment, warehouse, identity, email, carrier, batch | Owner confirms request and resulting state |
| Performance | Latency percentiles, throughput, errors, saturation under comparable load | Meets baseline-adjusted threshold |
| Resilience | Backup visibility, restore evidence, failover or component-loss test | Meets stated RTO/RPO evidence |
| Operations | Alarms, logs, traces, tickets, dashboards, on-call access and runbook | Operator detects and responds without migration team |
| Business | Test order through financially meaningful state | Business owner signs acceptance |
Record query text or test ID, input, expected output, actual output, timestamp, environment, executor, verifier, result, and artifact. A row count alone cannot detect wrong values; a sample alone cannot prove completeness. Use several techniques based on data criticality.
Prevent tests from creating real charges, shipments, emails, or customer notifications. Mark synthetic data and define cleanup.
12. Tool-state evidence without unsafe action
When the owner authorizes inspection, use narrow read-only calls and redact identifiers:
export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
aws mgn describe-source-servers \
--query 'items[].{LifeCycle:lifeCycle.state,DataReplication:dataReplicationInfo.dataReplicationState,Lag:dataReplicationInfo.lagDuration}'
aws dms describe-replication-tasks \
--query 'ReplicationTasks[].{Status:Status,MigrationType:MigrationType,Stats:ReplicationTaskStats}'
aws datasync list-tasks \
--query 'Tasks[].{Status:Status,Name:Name}'
The output is only one layer. MGN lifecycle state does not prove application correctness. DMS task status does not prove schema compatibility or zero discrepancies. A DataSync task status does not prove that the application can interpret every file or that the final delta is complete.
Capture relevant CloudWatch metrics, service events, logs, validation reports, and application evidence with UTC timestamps. Do not put account-specific commands into a general runbook without parameter validation and an authorization boundary.
13. Hypercare and operational transfer
Hypercare begins after traffic moves; it is not evidence that migration is complete. Define duration, coverage hours, enhanced dashboards, alert thresholds, incident severity, migration-team response, daily review, known defects, cost watch, and exit criteria.
Track:
- customer and business success rates;
- errors, latency, saturation, capacity and scaling;
- replication or queued-work residue;
- backup and restore-job outcomes;
- security findings and audit continuity;
- spend versus forecast; and
- support tickets by category and recurrence.
Operations acceptance requires resource inventory, ownership tags, diagrams, access, alarms, dashboards, backup/recovery procedure, patch and maintenance model, certificate/secret rotation, vendor support, service quotas, cost owner, known risks, and tested runbooks. The cloud operations owner signs the handoff; silence does not mean acceptance.
14. Decommission is a separate controlled change
Do not delete sources at cutover. First satisfy the agreed retention and rollback period, legal or records requirements, audit evidence retention, backup validation, license transfer, contract notice, data destruction method, CMDB and monitoring updates, DNS and certificate cleanup, account and credential revocation, invoice removal, asset disposal, and finance confirmation.
Quiesced source systems must be protected from accidental restart and unpatched exposure. Define network isolation, credential state, monitoring, retention cost, and authorized restoration. Decommission completion is proven by removed resources and stopped invoices, not by a ticket marked done.
15. Cost, risk, and retrospective
The wave budget includes target build, migration tooling, replication, snapshots and backups, data transfer, connectivity, temporary licenses, test environments, parallel source/target operation, staffing and overtime, support, hypercare, rollback reserve, retained source, and decommissioning. Map each to owner and active period.
After the wave, hold a blameless retrospective using actual timestamps and evidence. Compare planned versus actual duration, defects escaped from rehearsal, manual versus automated work, rollback margin, replication behavior, support load, cost, and business impact. Assign every improvement to the runbook template, platform, automation, discovery process, or training backlog before the next similar wave.
16. Capstone assignment
Submit p16-migration-wave-adr.md plus its referenced runbook and evidence register. Complete these stages:
- Reconstruct Orchid Orders' source architecture and identify unknown evidence.
- Define move groups and the transaction consistency boundary.
- Write measurable business, availability, RTO, RPO, security, and operational requirements.
- Compare three viable migration/cutover options and record the decision.
- Draw target identity, network, data, integration, failure, and observability paths.
- Build the five readiness-gate families with evidence and approvers.
- Calculate the latest rollback decision from tested durations.
- Create the countdown and minute-level runbook, including parallel work.
- Define GO/HOLD/FIX-FORWARD/ROLLBACK authority and triggers.
- Write separate rollback paths before and after target writes.
- Create the ten-layer validation matrix with exact tests and tolerances.
- Rehearse the plan as a tabletop with injected DNS cache, DMS latency, expired certificate, unavailable carrier contact, failed payment, and late business owner scenarios.
- Record decisions and update the runbook without erasing failed evidence.
- Define hypercare, operations acceptance, source retention, and decommission gates.
- Present a 15-minute architecture defense and a 10-minute command-center briefing.
The peer reviewer selects any runbook row and asks: What authorizes it? What proves success? What is the time limit? Who decides failure? What restores or compensates it? If any answer is missing, the package does not pass.
Diagnose a plan that looks complete
| Symptom | Hidden failure | Correction |
|---|---|---|
| All gates are green screenshots | Criteria, timestamps, and owners are absent | Build an evidence ledger with expiry and approval |
| Runbook says “team validates app” | No executable test or deadline | Name journey, input, expected result, owner, and duration |
| Rollback begins near outage end | Decision deadline ignored rollback duration | Schedule backward from validated rollback completion |
| DNS is the only traffic plan | Caches, TTL aging, sessions, and partner allowlists omitted | Model every client class and traffic mechanism |
| Database row counts match | Value, schema, sequence, LOB, and in-flight work can still differ | Combine reconciliation methods and exceptions |
| Target writes then rollback to old source | New transactions have no return path | Engineer replay, reverse flow, compensation, or fix-forward |
| Hypercare has no end | Operations acceptance criteria are absent | Define duration, exit metrics, and signed handoff |
| Source is “switched off” | Retention, security, contracts, and invoices continue | Create separate decommission evidence and approvals |
| One expert owns every critical task | Fatigue and key-person risk threaten execution | Add backups, rehearsal, separation, and shift limits |
Knowledge check
- What is the difference between an ADR and a runbook?
The ADR records why a decision is appropriate; the runbook controls exactly how an approved change executes.
- What is a move group?
The smallest dependency set that must move together to preserve behavior or consistency.
- How is the latest rollback decision calculated?
Subtract proven rollback, rollback validation, and contingency durations from the outage end.
- Why does rollback change after the first target write?
The source becomes stale and traffic reversal alone can lose accepted transactions.
- What makes a readiness gate objective?
Measurable criteria, timestamped evidence, owner, approver, expiry, and controlled waiver.
- Why is a healthy replication task insufficient?
It does not prove schema, values, application behavior, integrations, or business correctness.
- Who makes the go/no-go decision?
The named accountable authority within the documented deadline, informed by technical and business evidence.
- When may the source be deleted?
Only through a separate approved change after retention, recovery, records, security, licensing, and financial gates pass.
Lesson acceptance
The P16 capstone passes only when it contains:
- a signed, traceable ADR with alternatives and consequences;
- an exact scope, dependency graph, move groups, and consistency boundary;
- complete foundation, workload, data, test, people, and change gates;
- a backward-planned countdown and executable task-level runbook;
- explicit command authority and objective decision triggers;
- proven rollback timing and separate pre-write/post-write data handling;
- a ten-layer validation matrix with technical and business evidence;
- tabletop-injection results and controlled revisions;
- hypercare and operations acceptance criteria;
- separate retention and decommission controls;
- full transition and retained-source costs; and
- a peer review that can trace every material action to authorization, proof, deadline, owner, and recovery.