Lesson 336 · AWS Learning Path

AWS 336: Migration and disaster-recovery capstone

· Published · 8 min read

Labelled process diagram for AWS 336: Assessed business service to Migration wave and target architecture to Cutover plus protected recovery design to Business validation, DR test, failback, and decommission, with...

Why this capstone matters

Migration and disaster recovery cannot be separate assumptions. A workload moved to AWS must meet its recovery objectives in the target architecture, with current data, keys, identities, network, DNS, dependencies, runbooks, authority, and tests. Launching servers proves neither migration success nor recoverability.

This capstone joins assessment, strategy, landing-zone readiness, replication, cutover, post-write rollback, business validation, backup isolation, regional recovery, failover, reconciliation, failback, hypercare, decommissioning, and cost.

Scenario

Northstar's order service consists of 24 Linux VMs, a 4 TB PostgreSQL database changing at 12 MB/s peak, 8 TB shared files, corporate identity, payment and warehouse APIs, nightly finance batch, SMTP notifications, and on-premises DNS. It must leave a data center within nine months.

Business requirements:

  • checkout availability 99.95% monthly after migration;
  • migration cutover downtime below two hours;
  • normal-operation RPO below five minutes and RTO below one hour;
  • regional-disaster RPO below 15 minutes and RTO below four hours;
  • no duplicate charges or lost accepted orders;
  • order records retained seven years in approved geography;
  • source retained for rollback only until formally released;
  • first DR exercise within 30 days after migration.

Learning outcomes

By the end, you can build one executable migration/recovery design, distinguish MGN from DRS, prove data authority and rollback, measure business RTO/RPO, and present a funded roadmap with accepted residual risk.

1. Business impact and recovery objectives

Identify critical journeys, maximum tolerable disruption, financial/safety/legal impact, RTO, RPO, recovery priority, minimum viable service, and dependency objectives. RTO starts at disruption declaration and ends when the required business service is validated, not when an EC2 instance starts. RPO is measured against the latest accepted business data recovered.

JourneyMinimum recovery capabilityValidation evidence
Place orderIdentity, catalog, inventory, payment, order stateSynthetic order and reconciliation
RefundPayment and order history with authorizationApproved test transaction
Warehouse dispatchOrder export/event and partner pathEnd-to-end acknowledgment
Finance closeComplete, consistent period dataControl totals and sign-off

Resolve conflicting dependency objectives. A four-hour workload RTO is impossible if identity or DNS recovers in eight hours.

2. Current-state assessment

Inventory application/server/database/file components, versions, owners, data classes, dependencies, demand, maintenance, backup, incidents, licenses, certificates, DNS, firewall, batch, monitoring, and support. Observe a representative period including month-end and promotions.

Record facts, assumptions, unknowns, and confidence. Discover technical and organizational dependencies. Measure database change rate, file churn, network throughput, latency, and expected initial synchronization time. Include data growth and re-sync contingency.

3. Strategy by component

Apply retire, retain, rehost, relocate, repurchase, replatform, and refactor at component level. A defensible example may rehost application VMs with AWS Application Migration Service, replatform PostgreSQL after compatibility testing, transfer shared files with DataSync, and refactor notification later. It is not mandatory; compare alternatives.

Migration strategy and future DR strategy differ. MGN is designed for server migration with test/cutover lifecycle. AWS Elastic Disaster Recovery continuously protects supported source servers and includes recovery/failback capabilities. AWS documentation notes that MGN and DRS agents cannot be installed on the same source server simultaneously. Decide the transition rather than assuming MGN becomes DRS automatically.

4. Landing-zone and target readiness

Before replication, prove account/OU, identity, network/IP/DNS, hybrid bandwidth, egress/endpoints, security/logging, keys/secrets, backup, quotas, deployment, monitoring, support, cost ownership, and incident command.

Build target views for normal operation, migration coexistence, and regional recovery. Include multiple AZs, scaling, data, load balancing, sessions, downstream connectivity, batch, administration, logs, backups, and source/target authority.

5. Migration tool planes

Server replication

For MGN, document source prerequisites, agent identity, TCP 1500 path to staging replication servers, staging subnet/security, EBS/KMS, bandwidth, lag, launch settings, post-launch actions, test lifecycle, cutover, and finalization. Test boot, drivers, network, time, identity, application, monitoring, and performance.

Database migration

For DMS or native replication, document source logs/slots, full load plus CDC, schema conversion, keys, LOB handling, DDL behavior, target preparation, task capacity, validation, transaction consistency, lag, monitoring, and cleanup. Database replication readiness is not application cutover readiness.

Files

For DataSync, document agent/location authentication, network, include/exclude, metadata semantics, task mode/options, bandwidth, verification, final delta, ownership/permissions, and destructive options. Preserve a checksum/control manifest for critical records.

6. Capacity and schedule

Estimate transfer time:

minimum seconds = bytes x 8 / usable bits per second

Then include protocol/agent/storage bottlenecks, competing traffic, compression, change rate, retries, and validation. If source changes faster than effective replication, cutover lag never converges.

Model staging quotas, snapshots, volumes, DMS capacity, DataSync tasks, target instances, IPs, load balancers, database connections, and failover capacity. Reserve test windows and people, not only infrastructure.

7. Test migration

Launch isolated test instances without disrupting source replication. Use controlled DNS and synthetic identities/data. Validate functional journeys, integrations, performance, security, monitoring, backup/restore, deployment, batch, time zones, certificates, licenses, and operator runbooks.

Capture defects and repeat until exit criteria pass. Test HA/fault tolerance before cutover and complete a DR dry run after migration, consistent with the Migration Lens guidance.

8. Cutover state machine

  1. Change approval and communication confirmed.
  2. Source remains authoritative; target test evidence accepted.
  3. Freeze window begins where required.
  4. Final database/file/server synchronization and lag checks pass.
  5. Target infrastructure, keys, identity, network, telemetry, backup pass.
  6. Traffic shifts by defined DNS/load-balancer/client stages.
  7. Target writes begin; timestamp and transaction boundary recorded.
  8. Technical and business validation pass.
  9. Observation/hypercare starts.
  10. Migration finalization and source decommission require separate approval.

Define abort thresholds: lag, error, p95 latency, reconciliation variance, partner failure, security telemetry, and time remaining. Preserve a command/decision log.

9. Rollback after writes

Before target writes, rollback may restore source traffic after validating source currency. After target writes, “switch DNS back” can lose or duplicate orders. Define one of:

  • reverse replication of compatible changes;
  • transaction/event journal replay;
  • controlled dual-write with reconciliation;
  • compensating business operations;
  • forward repair in target if reversal is riskier.

Record source of truth at every stage, write fencing, idempotency, reconciliation keys, approval, and maximum rollback window. Test the method with synthetic transactions.

10. Target backup and DR architecture

Separate HA, backup, and regional DR. Multi-AZ handles common infrastructure faults; backup/PITR handles deletion/corruption/history; regional recovery addresses approved regional-disaster scenarios.

Compare backup/restore, pilot light, warm standby, and active-active against RTO/RPO, cost, complexity, data, and control-plane reliance. For this scenario, choose and defend one.

Protect backups with cross-account vault controls, separate identities, encryption keys, retention/lock where required, and restore testing. Replication alone can copy corruption or deletion.

11. Recovery sequence

Define declaration authority and recovery order, for example:

  1. incident scope, communication, and evidence;
  2. fence or confirm primary writes;
  3. establish identity, keys/secrets, DNS/network, and security visibility;
  4. recover authoritative database and validate RPO;
  5. recover files and application capacity;
  6. restore integrations, queues/batch, and notifications;
  7. shift traffic in stages;
  8. run technical and business acceptance;
  9. operate degraded/full service under incident command;
  10. plan reconciliation and failback separately.

Pre-provision whatever cannot fit inside RTO. Track control-plane calls required during an outage.

12. Failback

Failback is a planned migration to a selected primary, not the reverse of one toggle. Determine current source of truth, replication direction, divergence, new writes, reconciliation, capacity, deployment/config drift, DNS/traffic, observation, and approval. Do not fail back automatically when the old Region or data center returns.

13. Hypercare and decommission

Hypercare monitors user success, error/latency, replication/reconciliation, queues, integrations, batch, security, backups, cost, and incidents with explicit owners and exit criteria.

Decommission after business acceptance, rollback expiry, data retention/export, dependency proof, backup/recovery decision, DNS/certificate/identity removal, agent/replication cleanup, contract/license termination, CMDB update, and billing verification. Preserve audit evidence.

14. Read-only evidence track

aws sts get-caller-identity --query Arn --output text
aws mgn describe-source-servers --output table
aws dms describe-replication-tasks --output table
aws datasync list-tasks --output table
aws backup list-protected-resources --output table
aws drs describe-source-servers --output table

Empty or denied results describe current scope, not readiness. Do not install agents or start tests in this T0 capstone.

15. Capstone deliverables

Submit 24 artifacts:

  1. business impact and journey objectives;
  2. owner/decision/incident RACI;
  3. current inventory and evidence confidence;
  4. dependency map;
  5. component strategy decisions;
  6. target normal/coexistence/recovery views;
  7. landing-zone readiness gate;
  8. MGN/DRS/tool boundary ADR;
  9. database migration/validation plan;
  10. file transfer/control manifest plan;
  11. bandwidth/capacity/quota model;
  12. test-migration plan and exit evidence;
  13. cutover state machine/runbook;
  14. abort criteria;
  15. post-write rollback/forward-repair design;
  16. business reconciliation catalog;
  17. backup isolation and restore plan;
  18. DR strategy ADR and RTO/RPO budget;
  19. dependency-ordered recovery runbook;
  20. failover/fencing/traffic plan;
  21. failback/reconciliation plan;
  22. tabletop plus technical DR exercise;
  23. hypercare/decommission checklist;
  24. migration, coexistence, DR, transfer, support, and people cost model.

Failure injections

Test replication lag that never converges, target boot failure, database CDC loss, DNS clients pinned to source, payment timeout after charge, unavailable KMS key in recovery Region, and source recovery during target incident. For each record detection, decision authority, containment, data impact, recovery, and prevention.

Cost and cleanup

Model discovery, staging, replication, snapshots, transfer, test environments, target, coexistence, DR capacity, backup/copies, licenses, support, people, and delayed decommissioning. This design lesson creates no resources.

Knowledge check

  1. Does MGN automatically become DRS? No; they have different lifecycle and agent boundaries.
  2. When does RTO end? When the required business service is validated.
  3. Why is DNS-only rollback unsafe after writes? New authoritative transactions may be lost or duplicated.
  4. Is replication a backup? No; corruption/deletion can replicate.
  5. What is failback? A separately planned migration with authority and reconciliation.

Lesson acceptance

All 24 artifacts must form one executable plan. Data authority must be explicit at every stage, target HA/backup/DR must be tested, rollback must cover post-write behavior, business validation must define RTO/RPO completion, and decommissioning must end old risk and cost.

Official sources

Advertisement