Lesson 342 · AWS Learning Path

AWS 342: Architecture: complete the SAP migration and DR capstone

· Published · 7 min read

Labelled process diagram for AWS 342: Migration and BIA evidence to Reviewed target, wave, and recovery design to Tabletop decisions and corrected runbook to Business acceptance, drill cadence, and decommission, with...

Why this lesson matters

A migration is incomplete if the workload reaches AWS but cannot be operated or recovered. A disaster recovery design is incomplete if its RTO and RPO exist only in a slide. This final capstone defense joins discovery, migration waves, cutover, rollback, backup, regional recovery, failback, cost, and decommission into one evidence chain.

You will defend the migration and DR package from AWS336, inject failures, run a clocked tabletop, correct findings through AWS340, and produce an approval decision. The goal is to reason about business service continuity, not to memorize migration product names.

What you will be able to do

By the end, you can:

  • connect business impact, dependencies, migration strategy, target architecture, and recovery tier;
  • distinguish backup restoration, high availability, disaster recovery, and migration rollback;
  • define authoritative data and prevent split-brain decisions during cutover or failover;
  • measure RTO and RPO from evidence rather than repeat targets;
  • choose among rehost, relocate, replatform, refactor, repurchase, retain, and retire;
  • defend replication, orchestration, DNS, capacity, security, and failback choices;
  • identify when a migration wave or production cutover must not proceed.

Safety, scope, and inputs

This is a T0, document-based exercise. Do not start replication, launch test instances, alter DNS, or create recovery resources. Use sanitized evidence from approved environments or clearly labelled simulated evidence. Never include credentials, customer data, private IP inventories, or unredacted account IDs.

Bring all 24 artifacts from AWS336, the AWS340 findings register, business impact analysis, dependency inventory, wave plan, target diagrams, migration runbooks, cutover and rollback plan, recovery runbook, cost model, and approval evidence. If a required input is missing, record a finding and assess whether the exercise can safely continue.

One lifecycle, not two projects

Discover -> assess -> select strategy -> mobilize -> migrate wave
    |                                      |             |
    +-> business impact and dependencies   |             v
                                           |        cutover/rollback
                                           |             |
                                           v             v
                                     recovery design -> operate
                                           ^             |
                                           +-- drill/failback/decommission

Migration and recovery share dependency maps, data consistency rules, target capacity, network paths, identity, encryption, monitoring, and owners. Designing DR after cutover often discovers expensive assumptions too late.

Stage 1: evidence preflight

Create an index recording artifact, version, owner, evidence date, related application and wave, approval state, and sensitivity. Select one business service with at least three components, such as web tier, application tier, and database plus identity or messaging dependency.

Trace at least 20 requirements using this minimum schema:

FieldRequired answer
Business outcomeWhy migrate and what must not be harmed
Component/dependencyUpstream, downstream, owner, protocol, data direction
StrategyChosen migration approach and rejected alternative
TargetAccount, Region, AZ, network, compute, data, and security boundary
CutoverEntry criteria, sequence, authoritative data, validation
RollbackTrigger, deadline, data reconciliation, decision owner
RecoveryFailure scope, RTO, RPO, recovery pattern, failback
EvidenceDiscovery record, test result, metric, log, or approval
CostMigration, overlap, steady-state, DR, transfer, and retirement cost

Rate each as proven, partially proven, contradicted, or absent. Do not treat successful replication as proof of application correctness.

Stage 2: four defense panels

Panel A: assessment, strategy, and waves

Defend inventory completeness, dependency discovery, data classification, licensing, hardware constraints, business calendar, and application ownership. Explain why each component is rehosted, relocated, replatformed, refactored, repurchased, retained, or retired. State what new evidence would change that decision.

Wave ordering must respect dependencies and team capacity. Show pilot selection, factory throughput assumptions, coexistence, shared-service readiness, and learning feedback. Define stop conditions such as unknown critical dependencies, failed performance tests, unapproved downtime, or no rollback owner.

Panel B: target, tooling, security, and capacity

Trace source workload to migration tooling and target. Explain where AWS Application Migration Service, AWS Database Migration Service, DataSync, Snow Family, native database tools, or partner tooling fits, and do not force one tool onto every data type.

Defend landing-zone readiness, routing, DNS, certificates, secrets, IAM roles, encryption keys, logging, patching, backup, observability, and service quotas. Size for test, cutover overlap, normal demand, recovery demand, and growth. Identify temporary permissions and how they are removed.

Panel C: cutover, data authority, and rollback

Write the cutover as an ordered state machine: readiness approval, change freeze, final synchronization, source write handling, integrity checks, traffic change, business validation, observation, and closure. At each state define who may advance, pause, or reverse.

Name the authoritative write location at every moment. If rollback occurs after writes reach the target, explain how those writes return safely or why rollback is no longer allowed. Include queues, file transfers, scheduled jobs, external partners, caches, sessions, and TTL behavior. A DNS rollback alone does not reconcile data.

Panel D: availability, backup, DR, and failback

Separate four ideas: high availability handles local component faults; backup protects recoverable copies; DR restores service after a larger disruption; migration rollback returns from an unsuccessful change. One does not automatically provide the others.

For each service tier define failure scope, RTO, RPO, recovery pattern, replication lag, immutable or isolated recovery data, orchestration, dependency order, capacity, validation, communications, and failback. Explain pilot light, warm standby, active/passive, and multi-site tradeoffs. Prove backup restore and DR failover independently.

Stage 3: clocked tabletop

Use a facilitator, incident commander, technical owners, business decision owner, observer, and timekeeper. Start a visible clock at incident declaration. Record alert time, acknowledgement, declaration, containment, recovery start, data recovery point, service restoration, business validation, and failback decision.

Inject all seven events without warning:

  1. Available bandwidth falls and replication lag cannot meet the cutover window.
  2. A replicated server fails to boot in the target because of driver, boot mode, or network assumptions.
  3. Database change data capture stops and the last trustworthy transaction position is unclear.
  4. A payment is committed externally but the application times out during cutover.
  5. DNS caches split clients between source and target longer than planned.
  6. The recovery Region cannot use the required KMS key or secret.
  7. The primary Region returns while the recovery environment is accepting writes.

For each inject answer: what detects it, who decides, what evidence is trustworthy, where writes are permitted, whether to pause/rollback/fail over, how data is reconciled, and what customers are told. Record assumptions as findings.

Measuring recovery honestly

Target RTO is the approved maximum outage. Observed recovery time starts at the agreed event, usually detection or declaration, and ends only after technical and business validation. Report both when declaration delay is excluded.

Target RPO is the approved maximum data loss measured in time. Observed RPO is the difference between the incident’s last acknowledged valid write and the newest confirmed recovered write. Replication lag is useful evidence but is not automatically end-to-end RPO because buffers, queues, caches, and external systems may differ.

Record timeline timestamps, clock source, recovered transaction boundary, validation query, unresolved inconsistency, and confidence. If the tabletop cannot measure a value, label it unproven and schedule a controlled technical drill.

Critical fail conditions

The capstone cannot pass while any of these remains:

  • a critical dependency has no owner or migration disposition;
  • production cutover has no explicit entry, stop, rollback, or data-authority rule;
  • rollback assumes target writes can be discarded without business approval;
  • credentials, keys, quotas, DNS, or target capacity are untested critical assumptions;
  • RTO or RPO is claimed without a timestamped test method;
  • recovery omits identity, network, security, observability, or external dependencies;
  • failback has no consistency and split-brain prevention plan;
  • backups exist but restore evidence does not;
  • decommission can begin before retention, audit, reconciliation, and owner approval;
  • a critical finding is waived without an authorized risk owner and expiry.

Correction, approval, and decommission

Move every weakness through AWS340 with evidence, consequence, priority, owner, treatment, acceptance criteria, and retest. Update all dependent artifacts. Re-run affected tabletop stages and retain before-and-after evidence.

Approval may be rejected, conditional, or final. Conditional approval lists exact restrictions and deadlines. A cutover approval is wave-specific and time-bound; it is not permanent approval of the whole program.

Decommission only after business acceptance, observation period, data reconciliation, retention and legal checks, backup decision, monitoring transfer, CMDB and license updates, cost verification, and rollback expiry. Securely remove temporary migration access and stop duplicate charges.

Required submission artifacts

Submit these 18 items:

  1. final artifact index;
  2. business outcomes, BIA, RTO, and RPO summary;
  3. traceability matrix for at least 20 requirements;
  4. inventory and dependency confidence report;
  5. strategy decision register;
  6. wave plan with entry, exit, and stop criteria;
  7. target architecture and security data flows;
  8. tool selection and replication design;
  9. capacity, quota, and performance evidence;
  10. ordered cutover state machine;
  11. rollback and data-reconciliation decision tree;
  12. backup, DR, and failback runbook;
  13. cost model for migration, overlap, DR, and steady state;
  14. seven timestamped tabletop inject records;
  15. measured RTO/RPO worksheet and evidence limits;
  16. findings, corrections, and retest register;
  17. residual-risk, exception, and decommission register;
  18. stakeholder decision and oral-defense notes.

Lesson acceptance

All 18 artifacts must be internally consistent. Twenty requirements must trace to decisions and evidence. All seven failures must be reasoned through with named decision owners. Critical findings must be corrected and retested; observed and target RTO/RPO must never be presented as the same fact without proof.

Official sources

Advertisement