Lesson 292 · AWS Learning Path

AWS 292: Architecture: build a migration wave, cutover, rollback, and validation plan

· Published · 16 min read

Labelled process diagram for AWS 292: Approved assessment and target to Executable wave and runbook to Cutover or triggered rollback to Business acceptance, hypercare, and decommission evidence, with decision, proof...

Why this lesson matters

A migration can be technically possible and still fail during cutover because no one can answer four questions quickly: What happens next? What proves it worked? Who can decide? How do we return safely?

This capstone turns the outputs of P16 into an executable decision package for one business service. It is not another service overview. The learner must connect discovery, strategy, business case, target design, landing-zone readiness, replication, testing, change control, traffic movement, data ownership, rollback, business acceptance, hypercare, and decommissioning.

The result must be usable by people who did not write it, at an inconvenient hour, while facts change. Ambiguous verbs such as “validate,” “monitor,” and “rollback if required” are replaced with commands, expected results, deadlines, owners, and decision authority.

What you will be able to do

By the end, you can:

  • distinguish an architecture decision record, wave plan, cutover runbook, and evidence log;
  • prove that a dependency group is safe to move in one window;
  • build readiness gates that cannot be passed by opinion alone;
  • calculate the decision deadline from downtime and rollback duration;
  • define pre-write and post-write rollback branches;
  • sequence technical, business, communication, and change-management tasks;
  • validate infrastructure, security, data, integration, performance, recovery, and business outcomes;
  • run a command center with explicit authority and escalation;
  • transfer ownership through hypercare without silently decommissioning the source; and
  • defend the complete package in an architecture review.

Before you start

  • Use the supplied fictional service. Do not run a migration against a real workload.
  • Existing-account inspection requires owner approval and remains read-only. A source or target identifier can expose sensitive architecture; redact it from shared evidence.
  • Never stop a server, change DNS, disable a job, initiate DMS or MGN actions, execute a DataSync task, or alter routes while completing this lesson.
  • Do not place passwords, keys, tokens, private hostnames, customer records, or full account IDs in the runbook. Reference an approved secret and access process.
  • All times in the final runbook must use one explicit time zone and include the UTC equivalent.
  • A plan is not approved merely because a tool says replication is healthy. Named technical, business, security, operations, and change owners must accept their gates.

1. Know the four artifacts

ArtifactQuestion answeredChange frequencyApproval
Architecture decision record (ADR)Why this scope, target, strategy, and risk treatment were chosenWhen material assumptions or design changeArchitecture and accountable business/technology owners
Wave planWhich dependency groups move together, through which phases, and whenRolling planning cycleProgram and affected teams
Cutover runbookExactly who does what, when, using which evidence and recovery stepRehearsal findings and controlled updatesCutover/change authority
Execution and evidence logWhat actually occurred, result, timestamp, approver, and deviationDuring and after executionImmutable or controlled operational record

Do not hide operational detail inside the ADR. Do not turn a live runbook into a design essay. Link artifacts by stable version and change record so the command center knows which approved version is authoritative.

2. The supplied business service

Build the capstone for Orchid Orders, a fictional three-tier order service:

  • two Linux web servers behind an on-premises load balancer;
  • two Java application servers with in-memory sessions;
  • PostgreSQL primary plus reporting replica;
  • a file exchange used by warehouse and finance systems;
  • outbound payment, identity, email, and carrier APIs;
  • nightly settlement at 01:00 local time; and
  • 24x7 customer ordering with the lowest demand Saturday 22:00-00:00.

The approved direction is rehost web and application servers using AWS Application Migration Service, replatform PostgreSQL to Amazon RDS for PostgreSQL using AWS DMS full load plus change data capture, and move file exchange to an approved AWS storage target using DataSync. Traffic enters through the approved target load-balancing and DNS design. This is a teaching scenario, not a claim that these services fit every workload.

Requirements:

  • maximum approved customer outage: 90 minutes;
  • cutover window: 120 minutes, including the decision and rollback reserve;
  • RTO: 120 minutes; RPO at committed go-live: zero acknowledged orders lost;
  • target runs across two Availability Zones where the selected service supports it;
  • database and application tiers are private;
  • payment and warehouse integrations must pass before go-live;
  • a business owner must validate one test order through settlement-visible state;
  • source infrastructure must remain recoverable for 30 days; and
  • production source disposal requires separate retention, finance, security, and application-owner approval.

Known complications include fixed source IP allowlists, a 15-minute DNS TTL, in-memory sessions, one non-idempotent batch submission, a reporting user with an old PostgreSQL extension, and a carrier partner available only until 23:15.

3. Fix scope and the consistency boundary

A wave is a management grouping. A move group is the smallest set that must change together because splitting it would break business behavior or data consistency. One wave can contain multiple move groups with separate cutover decisions.

For every component record:

  • current and target identity;
  • application and infrastructure owner;
  • strategy and migration tool;
  • inbound and outbound dependencies;
  • data read/write behavior and system of record;
  • startup, shutdown, and health sequence;
  • maintenance window and blackout dates;
  • expected replication or transfer duration;
  • rollback method and maximum duration; and
  • source disposition.

Draw a dependency graph with direction, protocol, port, DNS name, identity, data classification, timeout, retry behavior, and owner. A CMDB relationship is a hypothesis until a representative observation window and an owner confirm it. Include scheduled, month-end, reporting, administrator, monitoring, backup, certificate, directory, and vendor flows that runtime discovery might miss.

Define the transactional consistency boundary. For Orchid Orders, order records, payment state, inventory request, and settlement event cannot be validated as unrelated servers. The plan must state where writes stop, how queues drain, how DMS latency reaches the accepted threshold, which database becomes authoritative, and how an accepted post-cutover order returns during rollback.

4. Write the ADR before the runbook

Use this ADR structure:

  1. Title, status, date, owners, reviewers, and superseded decisions.
  2. Decision statement in one paragraph.
  3. Business outcomes and measurable technical requirements.
  4. Scope, exclusions, move groups, and consistency boundary.
  5. Current architecture and evidence quality.
  6. Target architecture, accounts, Regions, Availability Zones, and trust boundaries.
  7. At least two viable alternatives.
  8. Selected strategy and migration pattern per component.
  9. Decision drivers, constraints, assumptions, and unknowns.
  10. Security, operations, resilience, data, and compliance consequences.
  11. Cost and schedule consequences, including parallel run.
  12. Risks, controls, owners, deadlines, and residual acceptance.
  13. Reversibility and point-of-no-simple-return.
  14. Validation, hypercare, decommission, and benefit-measurement approach.
  15. Approval, expiry, and revisit triggers.

Compare at least these options:

  • move the complete service in one window;
  • move read-only/reporting capability first and transactional capability later; and
  • retain the service temporarily while closing the IP, session, extension, and idempotency gaps.

Use weighted scoring only as decision support. A high score cannot override a mandatory security, licensing, RTO, or data-correctness requirement. Document why rejected options were viable and why they lost.

5. Convert readiness into evidence gates

Every gate needs an ID, criterion, evidence location, evidence time, owner, approver, status, expiry, and waiver authority. “Green” is not evidence.

Foundation gate

Prove target accounts and ownership, VPC/subnet capacity, routes, DNS, time synchronization, logging, security services, KMS key access, backup, quotas, monitoring destinations, support model, break-glass access, tagging, budgets, and configuration baselines. Verify both normal operator and emergency access before the window.

Workload gate

Prove target builds, immutable versions, patches, certificates, secrets, service identities, least privilege, security groups, dependency connectivity, storage, startup order, health checks, scaling, backup, restore, and operational runbooks. A successful ping does not prove TLS name, application authorization, transaction, or return path.

Data gate

Record initial-load completion, CDC state, source-log retention, replication latency trend, error tables, schema exceptions, LOB treatment, row or object reconciliation, checksums where meaningful, time zones, sequences, triggers, and target write protection. Document who can declare the final synchronization complete.

Test gate

Require functional, integration, security, performance, failover/recovery, monitoring, backup/restore, batch, operational, and business acceptance evidence. Compare target performance with a dated source baseline under comparable load. Record defects and accepted residual risk rather than deleting failed results.

People and change gate

Prove approved change, maintenance communication, participant availability, vendor contacts, access tests, bridge details, escalation tree, support tickets, freeze scope, duty limits, backup personnel, and go/no-go authority. Confirm that every named person accepted the role.

Waivers must identify the unmet criterion, reason, compensating control, risk owner, expiry, and approver. A waiver does not turn failed evidence into passed evidence.

6. Design backward from the deadline

The rollback decision must occur early enough to finish rollback before the approved outage ends.

latest rollback decision = outage end
                         - proven rollback duration
                         - rollback validation duration
                         - contingency reserve

For a 90-minute outage, if tested rollback takes 25 minutes, source validation takes 10 minutes, and reserve is 10 minutes, the latest decision is minute 45. A plan that starts business testing at minute 50 has already surrendered a safe rollback.

Build a countdown plan from at least T-30 days, T-14, T-7, T-3, T-1, T-4 hours, and T-30 minutes. Include TTL reduction early enough for caches to age, change approvals, final rehearsal, freeze, replication health, participant confirmation, support readiness, backups, and abort conditions. Restore normal TTL only after stability and rollback strategy allow it.

7. Build an executable runbook

Each task row must include:

FieldRequired content
ID and phaseStable reference and preparation, cutover, rollback, or recovery phase
Planned start/endAbsolute time plus expected duration
PredecessorTask or gate that must succeed first
Executor and verifierTwo roles where separation matters
Exact actionApproved command, console path, automation job, or communication text
Expected resultObservable value, state, count, or threshold
EvidenceTimestamped output or controlled link
Failure actionRetry once, investigate within limit, skip by waiver, stop, or roll back
Rollback counterpartTask restoring or compensating this change
ActualsStart, finish, result, deviation, and incident/change reference

Use explicit commands only after they have been tested in the same pattern and redacted for distribution. Avoid destructive wildcard operations and mutable “latest” artifacts. Every automation step needs inputs, version, dry-run or rehearsal evidence, idempotency behavior, timeout, output, and a manual recovery path.

Example timeline

21:30 bridge open, attendance and authority check
21:35 freeze verification and source health baseline
21:40 stop batch/schedulers; verify no new launches
21:45 place application in controlled write drain
21:50 drain sessions/queues; capture outstanding work
21:55 final file delta and database CDC catch-up
22:05 source writes stopped; consistency checkpoint recorded
22:10 final validation; GO-1 authorizes target write enablement
22:15 launch/enable target and perform technical smoke tests
22:25 shift controlled traffic and run integration tests
22:35 business order test and data reconciliation
22:45 final GO or rollback decision deadline
22:50 increase traffic only after GO
23:00 announce service restoration and begin hypercare

This example is incomplete until the learner supplies commands, owners, thresholds, rollback counterparts, and exact evidence. Parallel tasks must be visually identified; never assume row order alone conveys concurrency.

8. Establish command-center authority

Name these roles even if one person fills several:

  • cutover commander, who controls sequence and bridge discipline;
  • change authority, who confirms the approved window and material deviations;
  • technical migration lead;
  • application owner;
  • business decision owner;
  • source infrastructure, target cloud, network/DNS, database/data, security, and operations leads;
  • test coordinator and evidence recorder;
  • communications lead; and
  • incident manager if service impact exceeds the plan.

Only one role issues GO, HOLD, FIX FORWARD, and ROLLBACK calls. Technical teams recommend; the named accountable decision owner decides within the deadline. Predetermine what happens if that person is unreachable.

Use a concise cadence: task ID, expected result, actual result, variance, risk to deadline, recommendation, decision. Side conversations must return decisions to the main log. Never allow optimistic verbal status to replace evidence.

9. Define objective decision states

Use four states:

  • GO: all mandatory gates pass and remaining risk is accepted.
  • HOLD: no irreversible step; investigate within a fixed budget.
  • FIX FORWARD: target has writes or reverting is riskier, and repair can meet a defined deadline.
  • ROLLBACK: a trigger fired and enough time remains for the tested path.

Example rollback triggers:

  • mandatory payment or warehouse integration fails twice after one approved repair;
  • reconciliation exceeds the approved count or value tolerance;
  • sustained error rate, latency, or saturation exceeds a named threshold;
  • target database cannot accept writes safely;
  • security control or audit logging is absent;
  • business test cannot complete before the decision deadline; or
  • remaining rollback time falls below the proven requirement.

Avoid “major issue” or “performance unacceptable.” Those phrases transfer the decision to a debate under pressure.

10. Separate rollback before and after target writes

Before target writes

If the source remains authoritative and unchanged, rollback can stop target traffic, disable target processing, restore source schedulers and routes, return traffic, validate source operation, and announce restoration. Still account for DNS caches, sessions, queued messages, and external allowlists.

After target writes

Once orders are accepted in the target, the source is stale. Choose and rehearse one approach:

  • replicate target changes back to a prepared fail-forward source;
  • restore and replay an authoritative transaction or event log;
  • use a deliberately engineered dual-write pattern with reconciliation;
  • compensate business transactions under approved rules; or
  • remain on the target and fix forward.

Specify conflict handling, order preservation, sequence values, duplicate prevention, encryption, transfer time, validation, and the owner who accepts any loss boundary. Native backup and restore is useful only when tested restore time and data point satisfy the deadline.

Define the point-of-no-simple-return. This can be first target write, irreversible schema change, partner allowlist removal, destructive source change, or replication finalization. Require a specific authorization immediately before it.

11. Build a validation matrix

Validation must test a chain, not isolated green dashboards.

LayerExample evidenceAcceptance
InfrastructureTarget state, AZ placement, capacity, quota, volume and load balancer healthMatches approved design and headroom
Network/DNSResolution from each client class, TLS, forward and return paths, partner source IPEvery required flow succeeds; forbidden flow fails
Identity/securityWorkload identity, least privilege, secret retrieval, encryption, audit eventsAllowed and denied cases behave as designed
DataCounts, control totals, sampled records, sequence/encoding/time-zone checks, DMS exceptionsWithin signed tolerance with no unexplained error
ApplicationLogin, create/read/update, session, file, API, and error behaviorCritical journeys pass
IntegrationPayment, warehouse, identity, email, carrier, batchOwner confirms request and resulting state
PerformanceLatency percentiles, throughput, errors, saturation under comparable loadMeets baseline-adjusted threshold
ResilienceBackup visibility, restore evidence, failover or component-loss testMeets stated RTO/RPO evidence
OperationsAlarms, logs, traces, tickets, dashboards, on-call access and runbookOperator detects and responds without migration team
BusinessTest order through financially meaningful stateBusiness owner signs acceptance

Record query text or test ID, input, expected output, actual output, timestamp, environment, executor, verifier, result, and artifact. A row count alone cannot detect wrong values; a sample alone cannot prove completeness. Use several techniques based on data criticality.

Prevent tests from creating real charges, shipments, emails, or customer notifications. Mark synthetic data and define cleanup.

12. Tool-state evidence without unsafe action

When the owner authorizes inspection, use narrow read-only calls and redact identifiers:

export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
aws mgn describe-source-servers \
  --query 'items[].{LifeCycle:lifeCycle.state,DataReplication:dataReplicationInfo.dataReplicationState,Lag:dataReplicationInfo.lagDuration}'
aws dms describe-replication-tasks \
  --query 'ReplicationTasks[].{Status:Status,MigrationType:MigrationType,Stats:ReplicationTaskStats}'
aws datasync list-tasks \
  --query 'Tasks[].{Status:Status,Name:Name}'

The output is only one layer. MGN lifecycle state does not prove application correctness. DMS task status does not prove schema compatibility or zero discrepancies. A DataSync task status does not prove that the application can interpret every file or that the final delta is complete.

Capture relevant CloudWatch metrics, service events, logs, validation reports, and application evidence with UTC timestamps. Do not put account-specific commands into a general runbook without parameter validation and an authorization boundary.

13. Hypercare and operational transfer

Hypercare begins after traffic moves; it is not evidence that migration is complete. Define duration, coverage hours, enhanced dashboards, alert thresholds, incident severity, migration-team response, daily review, known defects, cost watch, and exit criteria.

Track:

  • customer and business success rates;
  • errors, latency, saturation, capacity and scaling;
  • replication or queued-work residue;
  • backup and restore-job outcomes;
  • security findings and audit continuity;
  • spend versus forecast; and
  • support tickets by category and recurrence.

Operations acceptance requires resource inventory, ownership tags, diagrams, access, alarms, dashboards, backup/recovery procedure, patch and maintenance model, certificate/secret rotation, vendor support, service quotas, cost owner, known risks, and tested runbooks. The cloud operations owner signs the handoff; silence does not mean acceptance.

14. Decommission is a separate controlled change

Do not delete sources at cutover. First satisfy the agreed retention and rollback period, legal or records requirements, audit evidence retention, backup validation, license transfer, contract notice, data destruction method, CMDB and monitoring updates, DNS and certificate cleanup, account and credential revocation, invoice removal, asset disposal, and finance confirmation.

Quiesced source systems must be protected from accidental restart and unpatched exposure. Define network isolation, credential state, monitoring, retention cost, and authorized restoration. Decommission completion is proven by removed resources and stopped invoices, not by a ticket marked done.

15. Cost, risk, and retrospective

The wave budget includes target build, migration tooling, replication, snapshots and backups, data transfer, connectivity, temporary licenses, test environments, parallel source/target operation, staffing and overtime, support, hypercare, rollback reserve, retained source, and decommissioning. Map each to owner and active period.

After the wave, hold a blameless retrospective using actual timestamps and evidence. Compare planned versus actual duration, defects escaped from rehearsal, manual versus automated work, rollback margin, replication behavior, support load, cost, and business impact. Assign every improvement to the runbook template, platform, automation, discovery process, or training backlog before the next similar wave.

16. Capstone assignment

Submit p16-migration-wave-adr.md plus its referenced runbook and evidence register. Complete these stages:

  1. Reconstruct Orchid Orders' source architecture and identify unknown evidence.
  2. Define move groups and the transaction consistency boundary.
  3. Write measurable business, availability, RTO, RPO, security, and operational requirements.
  4. Compare three viable migration/cutover options and record the decision.
  5. Draw target identity, network, data, integration, failure, and observability paths.
  6. Build the five readiness-gate families with evidence and approvers.
  7. Calculate the latest rollback decision from tested durations.
  8. Create the countdown and minute-level runbook, including parallel work.
  9. Define GO/HOLD/FIX-FORWARD/ROLLBACK authority and triggers.
  10. Write separate rollback paths before and after target writes.
  11. Create the ten-layer validation matrix with exact tests and tolerances.
  12. Rehearse the plan as a tabletop with injected DNS cache, DMS latency, expired certificate, unavailable carrier contact, failed payment, and late business owner scenarios.
  13. Record decisions and update the runbook without erasing failed evidence.
  14. Define hypercare, operations acceptance, source retention, and decommission gates.
  15. Present a 15-minute architecture defense and a 10-minute command-center briefing.

The peer reviewer selects any runbook row and asks: What authorizes it? What proves success? What is the time limit? Who decides failure? What restores or compensates it? If any answer is missing, the package does not pass.

Diagnose a plan that looks complete

SymptomHidden failureCorrection
All gates are green screenshotsCriteria, timestamps, and owners are absentBuild an evidence ledger with expiry and approval
Runbook says “team validates app”No executable test or deadlineName journey, input, expected result, owner, and duration
Rollback begins near outage endDecision deadline ignored rollback durationSchedule backward from validated rollback completion
DNS is the only traffic planCaches, TTL aging, sessions, and partner allowlists omittedModel every client class and traffic mechanism
Database row counts matchValue, schema, sequence, LOB, and in-flight work can still differCombine reconciliation methods and exceptions
Target writes then rollback to old sourceNew transactions have no return pathEngineer replay, reverse flow, compensation, or fix-forward
Hypercare has no endOperations acceptance criteria are absentDefine duration, exit metrics, and signed handoff
Source is “switched off”Retention, security, contracts, and invoices continueCreate separate decommission evidence and approvals
One expert owns every critical taskFatigue and key-person risk threaten executionAdd backups, rehearsal, separation, and shift limits

Knowledge check

  1. What is the difference between an ADR and a runbook?

The ADR records why a decision is appropriate; the runbook controls exactly how an approved change executes.

  1. What is a move group?

The smallest dependency set that must move together to preserve behavior or consistency.

  1. How is the latest rollback decision calculated?

Subtract proven rollback, rollback validation, and contingency durations from the outage end.

  1. Why does rollback change after the first target write?

The source becomes stale and traffic reversal alone can lose accepted transactions.

  1. What makes a readiness gate objective?

Measurable criteria, timestamped evidence, owner, approver, expiry, and controlled waiver.

  1. Why is a healthy replication task insufficient?

It does not prove schema, values, application behavior, integrations, or business correctness.

  1. Who makes the go/no-go decision?

The named accountable authority within the documented deadline, informed by technical and business evidence.

  1. When may the source be deleted?

Only through a separate approved change after retention, recovery, records, security, licensing, and financial gates pass.

Lesson acceptance

The P16 capstone passes only when it contains:

  • a signed, traceable ADR with alternatives and consequences;
  • an exact scope, dependency graph, move groups, and consistency boundary;
  • complete foundation, workload, data, test, people, and change gates;
  • a backward-planned countdown and executable task-level runbook;
  • explicit command authority and objective decision triggers;
  • proven rollback timing and separate pre-write/post-write data handling;
  • a ten-layer validation matrix with technical and business evidence;
  • tabletop-injection results and controlled revisions;
  • hypercare and operations acceptance criteria;
  • separate retention and decommission controls;
  • full transition and retained-source costs; and
  • a peer review that can trace every material action to authorization, proof, deadline, owner, and recovery.

Official sources

Advertisement