AWS 344: Continuous improvement and migration
Why this checkpoint matters
Architects spend more time improving existing systems than designing empty environments. Brownfield work begins with incomplete evidence, coupled releases, undocumented dependencies, operational pain, and business deadlines. Migration adds temporary coexistence, data authority, rollback, and decommission risk.
This checkpoint tests whether you can diagnose before prescribing, choose different migration strategies for different components, learn between waves, and prove that an improvement changed a measurable outcome.
Outcomes
You will demonstrate that you can:
- reconstruct current behavior from architecture, metrics, logs, traces, incidents, costs, and interviews;
- separate symptom, proximate cause, contributing condition, and organizational cause;
- prioritize improvements by impact, urgency, effort, dependency, and reversibility;
- assign the seven migration strategies at component level;
- design waves, cutover, rollback, recovery, hypercare, and decommissioning;
- turn each wave into evidence for the next wave;
- measure outcome, guardrail, operational, financial, and recovery indicators.
Rules and supplied case
This is a T0 document exercise. Do not change production or create AWS resources. Label simulated evidence. Record unknowns instead of silently filling them.
Falcon Retail operates a ten-year-old order platform on premises:
- two web servers and four Java application servers behind a hardware load balancer;
- Oracle database, NFS document share, RabbitMQ, nightly SFTP exports, LDAP, and SMTP relay;
- 99.9-percent stated availability, but no approved RTO/RPO;
- p95 checkout latency rises from 800 ms to 8 seconds during promotions;
- application CPU remains below 40 percent while database sessions and NFS latency rise;
- deployments require six hours and failed changes often restore VM snapshots;
- one payment API, one warehouse VPN, and two undocumented batch consumers;
- 120-day log retention and seven-year invoice retention;
- infrastructure allocation costs USD 92,000 per month, but shared-license costs are missing;
- the data center contract ends in nine months.
Leadership says, “Move everything quickly, modernize it, cut cost 35 percent, and never take downtime.” Convert that conflict into an executable program.
Part 1: establish the evidence baseline
Build an evidence register with source, time window, owner, confidence, scope, and limitation. Require at least:
- traffic rate, concurrency, latency percentiles, error rate, and user-journey completion;
- host CPU, memory pressure, disk latency/queue, network loss, and process saturation;
- database waits, sessions, locks, query latency, transaction rate, growth, and restore history;
- queue depth, oldest-message age, retries, dead letters, and consumer throughput;
- file count/size/change rate, metadata needs, and NFS dependency inventory;
- change failure rate, deployment duration, rollback evidence, incident MTTA/MTTR;
- application, network, identity, licensing, and organizational dependencies;
- invoices, allocation gaps, contracts, commitments, and decommission liabilities.
Draw current request, identity, data, network, deployment, and failure paths. Mark every dependency observed, reported, inferred, or unknown. One day of flow logs cannot prove that a monthly batch dependency does not exist.
Part 2: diagnose before designing
For promotion latency, construct at least four competing hypotheses: database lock contention, NFS serialization, synchronous payment calls, connection-pool exhaustion, or application garbage collection. For each state supporting evidence, contradicting evidence, missing evidence, and the safest discriminating test.
customer symptom -> service indicator -> saturated dependency
-> causal mechanism -> contributing condition
-> control/process weakness -> measurable correction
Low average CPU does not prove spare capacity. A workload can wait on database, storage, locks, network, quotas, or external calls. Resizing compute is not accepted unless evidence connects compute saturation to the customer symptom.
Create a baseline with outcome metric, current distribution, source, target, guardrail, owner, and review window. Include availability, p95/p99 latency, order correctness, deployment safety, recovery, cost per order, and operator toil.
Part 3: improvement portfolio
Identify at least 15 findings across all six Well-Architected pillars. Process each through AWS340. Divide action into:
- stabilize now: prevent loss, exposure, or repeated severe incidents;
- prepare: improve inventory, observability, tests, automation, landing zone, and skills;
- migrate: move with a defined strategy and measurable gate;
- modernize: change architecture where value exceeds transition risk;
- retire/decommission: remove cost and attack surface after evidence permits.
Score actions by impact, risk reduction, urgency, effort, prerequisite, and reversibility. Do not put a large refactor on the critical path merely because it is attractive. Include a 30/60/90-day plan and nine-month roadmap.
Part 4: component strategies and target decisions
Assign retire, retain, rehost, relocate, repurchase, replatform, or refactor to web, Java, Oracle, NFS, RabbitMQ, SFTP, LDAP, SMTP, reporting consumers, and each external integration. State evidence, rejected alternatives, reversibility, owner, and reevaluation trigger.
Compare these boundaries:
- EC2 rehost versus containers versus managed/serverless compute;
- Oracle on EC2/RDS/custom versus compatible engine conversion;
- EFS/FSx/S3 based on POSIX, locking, latency, object semantics, and client change;
- Amazon MQ versus SQS/SNS/EventBridge/MSK based on protocol and delivery semantics;
- Transfer Family/DataSync/native transfer based on partner protocol and migration need;
- Managed Microsoft AD, AD Connector, IAM Identity Center, or application identity change;
- synchronous integration versus queue/event decoupling with idempotency and reconciliation.
The target must preserve correctness during coexistence. Define source of truth, write authority, replication direction, schema compatibility, identity, routing, monitoring, and support ownership for every transition state.
Part 5: waves and cutover
Build at least three waves: pilot, representative business wave, and critical order-platform wave. A wave is a dependency-aware group with shared preparation and bounded cutover, not merely a calendar list.
Each wave needs entry criteria, scope, owner, test environment, performance baseline, security and data validation, cutover sequence, stop conditions, rollback deadline, post-write treatment, business acceptance, hypercare exit, and decommission gate.
Write the order cutover as a state machine:
- Readiness and authority confirmed.
- Change freeze and final synchronization started.
- Source writes stopped, queued, or explicitly dual-written.
- Consistency boundary recorded.
- Routing changed.
- Technical and business transactions validated.
- Observation window entered.
- Rollback remains safe or is formally closed.
- Source enters retention, then approved decommission.
Explain what happens to orders accepted after target activation. A rollback that restores traffic but loses those writes is not a rollback.
Part 6: recovery and failure injects
Define workload-specific RTO and RPO. Separate Multi-AZ high availability, backup/restore, disaster recovery, and migration rollback. Complete eight tabletop injects:
- Dependency discovery missed a monthly reporting consumer.
- Replication lag exceeds the cutover window.
- The target boots but cannot reach LDAP or DNS.
- DMS full load completes while CDC reports errors on keyless tables.
- Payment commits but the application times out.
- DNS sends clients to both environments beyond the planned interval.
- Recovery cannot decrypt data because key permissions differ.
- Source decommission starts before audit retention is confirmed.
Record detection, decision owner, pause/rollback/failover choice, write authority, reconciliation, customer communication, elapsed time, and permanent improvement.
Part 7: continuous improvement loop
After every wave compare target and observed results. Run a Well-Architected review on the migrated state; update seven-strategy decisions, wave plan, estimates, tests, runbooks, and training. Track disproved hypotheses as well as lessons learned.
| Category | Examples |
|---|---|
| Outcome | Checkout completion, correctness, availability |
| Performance | p95/p99 latency, saturation, backlog age |
| Delivery | Lead time, deployment duration, change failure rate |
| Recovery | Observed restore/failover time, recovered transaction boundary |
| Cost | Cost per order, overlap burn, forecast variance, retired spend |
| Operations | Pages, toil hours, stale runbooks, owner coverage |
| Guardrail | Security findings, unauthorized access, residency breach |
An improvement passes only when its acceptance window and guardrails pass. If latency improves while correctness or cost breaches a limit, the change is not accepted.
Scoring and pass gate
Score 100 points: evidence/baseline 15, diagnosis 15, improvement portfolio 15, strategy/target 15, waves/cutover 15, recovery/failure reasoning 15, measurement/governance 10. Pass requires 80 overall and at least 50 percent in every category.
Automatic failure: no authoritative-write rule, rollback discards accepted writes, RTO/RPO presented as measured without evidence, critical dependency has no owner, decommission precedes retention approval, or savings omit migration overlap and operating ownership.
Required artifacts
Submit an evidence register, six current-state views, hypothesis worksheet, baseline scorecard, 15-finding register, prioritized backlog, 30/60/90-day plan, component strategy matrix, target/coexistence views, three wave dossiers, cutover state machine, recovery design, eight inject records, cost/TCO scenarios, post-wave review, corrected roadmap, residual-risk register, and decision approval.