Lesson 294 · AWS Learning Path

AWS 294: AWS Elastic Disaster Recovery

· Published · 16 min read

Labelled process diagram for AWS 294: Protected source and agent to Continuous block replication to staging to Drill or recovery launch to Dependency validation, failback, and cleanup, with decision, proof and...

Why this lesson matters

AWS Elastic Disaster Recovery (AWS DRS) continuously replicates supported source-server disks into a low-cost staging area and can launch recovery instances from selected points in time. That is powerful, but disk replication is not application recovery.

A green replication state does not prove that three dependent servers start in the right order, the database is transactionally usable, DNS reaches the recovery environment, identity works, licenses permit the launch, operators can access the system, or the business accepts the result. A recent recovery point can also contain the same corruption or malicious change that caused the disaster.

This lesson teaches the whole lifecycle: suitability, source agent, network, staging resources, recovery points, launch configuration, isolated drills, disaster declaration, traffic movement, RPO/RTO evidence, reverse replication, failback, and cleanup. The architect's job is to transform block-level recovery capability into a tested service-recovery plan.

What you will be able to do

By the end, you can:

  • decide when DRS fits better than backup restore, application-native recovery, or database replication;
  • trace control and block-replication traffic from a source to the staging area;
  • explain source servers, replication servers, staging EBS volumes, snapshots, conversion, and recovery instances;
  • calculate bandwidth, replication lag, point age, and complete service RPO/RTO;
  • design least-privilege, private-network, encryption, and evidence controls;
  • configure conceptually distinct replication, DRS launch, EC2 launch-template, and post-launch settings;
  • select a recovery point without copying corruption blindly;
  • design isolated, nondisruptive recovery drills for dependency groups;
  • run failover and failback with one data authority; and
  • diagnose replication, launch, boot, network, application, and reverse-replication failures.

Before you start

  • Use supplied evidence. Do not install an agent, initialize DRS, start a drill or recovery, reverse replication, disconnect a source, or terminate an instance.
  • DRS actions can create EC2, EBS, snapshot, networking, log, and data-transfer cost. Existing-account inspection requires owner approval and is read-only.
  • Never expose installer credentials, account IDs, source names, private IPs, tags, recovery addresses, logs with customer data, or post-launch secrets.
  • Verify current supported operating systems, Regions, quotas, licensing rules, and agent prerequisites for every real design.
  • Coordinate with backup, database, security, network, licensing, application, and business owners. DRS does not replace those responsibilities.

1. Choose the recovery mechanism from requirements

Requirement or workloadDRS directionAlternative to evaluate
Full supported server, low RPO, and rapid EC2 launchStrong candidateApplication-native replication where it gives better consistency
Database with strict transaction consistency and near-zero RPOProtect host only as one layerNative database replication, managed multi-Region database, logs and backups
Accidental deletion or corruption discovered days laterPoint-in-time range might help, but contamination can replicateIndependent immutable backup and longer retention
SaaS or managed AWS serviceDRS cannot install an agent into the managed hostService-native backup, replication, export, or reconstruction
Kubernetes or container platformDRS replicates whole supported servers, not logical cluster stateImages, manifests, GitOps, persistent-data protection, application recovery
Unsupported OS, storage, boot model, or driverReject until support is provedBackup restore, replatform, native tooling, or another product
One application spanning many serversDRS can recover serversOrchestration and application-consistent sequencing remain your responsibility

DRS is distinct from AWS Application Migration Service. Both use related continuous replication concepts, but DRS is operated as an ongoing disaster-recovery capability with drills, recovery, and failback. MGN is oriented toward migration testing, cutover, and finalization. Select the service based on the operating lifecycle, not a familiar console.

Keep an independent backup strategy. Continuous replication can copy encryption events, logical corruption, malware, or human mistakes. Recovery-point history has a service-defined window and is not a substitute for long-term, immutable, policy-governed backups.

2. Understand the architecture

supported source server
  |-- AWS Replication Agent -- TCP 443 --> DRS Regional service endpoint
  |
  +-- changed disk blocks -- TCP 1500/TLS --> replication server
                                              |
                                              v
                                     staging EBS volumes
                                              |
                                  snapshots / recovery points
                                              |
                          drill or recovery launch and conversion
                                              |
                                              v
                                      recovery EC2 instance

Source server and agent

The AWS Replication Agent runs on each protected supported source server. It observes block changes on selected disks and sends them continuously. Confirm OS/kernel support, free space and memory prerequisites, root/administrator installation authority, disk layout, boot mode, proxy behavior, endpoint access, anti-malware exceptions, and performance impact before deployment.

The agent sees blocks, not business transactions. It does not know that an application transaction spans a database, queue, file server, and API. A recovery point is generally crash-consistent at the protected server/disk level unless the application plan creates stronger consistency through quiescing, native tooling, or coordinated recovery.

Staging area

DRS launches and manages lightweight replication servers in a staging subnet and creates staging EBS volumes corresponding to protected source volumes. Staging is not the production recovery environment. It is the continuously updated replication substrate.

Choose the staging VPC, subnet, routes, service/S3 access, security groups, private or public replication path, instance types, EBS choices, encryption key, resource tags, and bandwidth throttling deliberately. Staging must survive the failure scenario it protects against. Replicating into the same failed Region or an account that operators cannot access does not meet a Regional/account-isolation objective.

DRS manages replication-server lifecycle, so do not treat those instances as ordinary application servers or manually patch and customize them. Monitor service state and the configuration boundary instead.

Recovery resources

During a drill or recovery, DRS uses selected recovery-point data and launch settings to create converted, bootable EC2 recovery instances. These are separate from staging replication servers and volumes. Recovery instances need the full target VPC, routing, security, identity, DNS, certificates, secrets, load balancing, monitoring, backups, licenses, and operational ownership required by the application.

3. Prove the network and bandwidth paths

The source agent communicates with the Regional DRS service over TCP 443. It sends replicated data to replication servers in the staging subnet over TCP 1500. Staging resources also need required TCP 443 access to DRS, EC2, S3 service resources, and current package repositories documented by AWS. If VPC endpoints are used, route, DNS, security, and endpoint policies must permit every required operation and bucket.

TCP 1500 replication is encrypted with TLS. Encryption in transit does not eliminate the need for restricted routes, source CIDRs, security-group review, network monitoring, and encrypted EBS volumes. The service-managed security group can maintain required replication rules; understand that behavior before enforcing a competing configuration-remediation rule.

Draw both directions and all stateful/stateless controls:

  • source route and corporate firewall;
  • VPN, Direct Connect, internet, or other approved transport;
  • destination route table and subnet;
  • security group and network ACL return ports;
  • DRS, EC2, and S3 service access;
  • DNS and proxy path; and
  • failback path from recovery instance toward the target source environment.

Bandwidth and lag math

Sustained available replication throughput must exceed the source's average changed-block rate. Otherwise lag grows.

net catch-up rate = effective replication throughput - source change rate
catch-up time      = queued changed data / net catch-up rate

If sources change 160 Mbps and the effective path delivers 200 Mbps, only 40 Mbps remains for catching up. A 180 GB backlog at 40 Mbps takes roughly 10 hours before protocol and workload variation. Calculate in consistent bits/bytes and test representative peaks.

Initial synchronization must transfer allocated/protected disk blocks as the service requires, not merely application file sizes. Rescans after disk or unexpected-state changes can raise demand. Throttling protects production links but can violate RPO if it is below the write rate.

4. Design identity, encryption, and governance

Separate these identities:

  • administrator who initializes and configures DRS;
  • temporary or controlled installer authorization;
  • service-linked and service roles used by DRS;
  • EC2 instance profile used by launched recovery instances;
  • post-launch action role and Systems Manager access; and
  • operators allowed to start drills, start recovery, reverse replication, or terminate resources.

Do not store permanent AWS access keys on source servers. Follow the current documented installer authentication method and remove temporary privilege after enrollment. Scope human roles by account, Region, action, tags, and change process where supported. A read-only observer must not also be able to initiate recovery or delete source-server records.

Use an approved KMS key strategy for staging and recovery EBS volumes and snapshots. Validate that DRS/service roles, recovery launch paths, copied snapshots, cross-account patterns, and operators retain access during a disaster. A customer-managed key whose policy depends on the failed environment can block recovery.

Record DRS API actions in CloudTrail, monitor configuration drift, protect log integrity, tag resources, and alert on unexpected recovery launches, agent disconnection, lag, stalled sync, and setting changes. Define emergency access that works when the normal identity provider or network is unavailable, then test and audit it.

5. Keep the setting layers separate

Replication configuration

This controls how source data reaches staging: staging subnet, replication-server choices, routing/public-IP behavior where applicable, security groups, EBS type, encryption, bandwidth throttling, and automatic handling of new disks. It affects RPO and steady cost.

DRS launch settings

These govern DRS recovery behavior, such as instance-type right-sizing, start-on-launch, licensing choices, copy-private-IP behavior where supported, tags, and recovery mode. Review rather than accepting right-sizing blindly; a drill must prove performance and licensing.

EC2 launch template

The launch template provides target EC2 details such as VPC/subnet, security groups, instance type overrides, IAM instance profile, disks, key or access design, and other EC2 properties. It is a versioned dependency. Prove which version DRS will use and control changes to it.

Post-launch actions

Post-launch actions can install or configure agents, run automation, and perform verification. Make them versioned, idempotent, timed, observable, and safe to retry. Separate mandatory recovery actions from optional improvements. A failed optional monitoring installation should not necessarily block boot; a failed security or application configuration might.

Never place secrets directly in scripts or user data. Retrieve them through workload identity from an approved store. Capture action result and logs as evidence.

6. Interpret lifecycle and replication evidence

For each source, track:

  • source-server lifecycle and last-seen time;
  • initial-sync or rescan progress;
  • data replication state;
  • lag duration and trend, not one sample;
  • stalled disks and error messages;
  • last recovery-point time and available point range;
  • launch-setting review status;
  • last drill date and result; and
  • membership in an application recovery group.

“Continuous Data Replication” means the replication pipeline is operating; it does not prove zero lag or recoverability. “Ready” cannot mean only that DRS allows a launch. Create an application readiness status that requires an accepted drill, current runbook, target dependencies, recovery point, access, capacity, data validation, and owner sign-off.

Measure effective RPO from the newest application-usable recovery point, not from the current clock to an attractive dashboard timestamp. If replication lag is 4 minutes but recovery must use a 25-minute-old point before corruption, effective data loss exposure is approximately 25 minutes plus any application consistency gap.

7. Select a recovery point safely

Latest is appropriate for infrastructure failure when the newest replicated state is trustworthy. It can be the worst choice after ransomware, file corruption, bad deployment, or destructive administration.

Build a timeline from:

  • first user symptom;
  • first malicious or corrupt write;
  • monitoring and audit events;
  • backup and snapshot evidence;
  • replication lag; and
  • last known-good business transaction.

Launch candidate points in an isolated network. Run malware/security analysis, filesystem and database recovery, application startup, data reconciliation, and business sampling. Do not connect two candidates to production dependencies or allow both to process real work.

For multi-server applications, independent block points might not represent one coordinated business instant. Define startup order and use application-native logs, replication, transaction recovery, queue replay, or reconciliation to reach a consistent service state.

8. Engineer an isolated recovery drill

A recovery drill uses the same source launch settings and point-in-time mechanisms as a real recovery and should not disrupt ongoing source replication. The source's new changes continue to staging, not into an already launched drill instance.

Isolation prevents duplicate production behavior. Use a drill VPC/subnets or rigorously controlled security groups/routes, nonproduction DNS, blocked outbound side effects, disabled schedulers, safe identities, synthetic data, and service virtualization where needed. Prevent payment, email, shipment, directory writes, license conflicts, and source/target split brain.

The drill must test:

  1. disaster declaration and emergency access;
  2. recovery-point selection;
  3. dependency-group launch order and conversion;
  4. boot, OS, disk, driver, time, and hostname state;
  5. identity, DNS, routing, security, certificates, and secrets;
  6. database/file/application consistency;
  7. post-launch actions and observability;
  8. load, capacity, quotas, and scaling;
  9. backup and restore in the recovery environment;
  10. traffic-switch procedure without affecting production;
  11. operator and business validation; and
  12. cleanup and proof that source replication remains healthy.

Record timestamps for declaration, launch request, each instance ready, application ready, business accepted, and simulated traffic restoration. RTO ends when the required business service is usable and accepted, not when EC2 reports running.

9. Run failover as a controlled state machine

Use these states:

protected -> incident suspected -> disaster declared
          -> source writes fenced -> point selected
          -> recovery launched -> technical validation
          -> business validation -> traffic switched
          -> operating in recovery -> failback prepared

Before launch, confirm disaster scope, point choice, target account/Region access, quotas, capacity, launch template, dependencies, and change authority. Fence source writes or prevent simultaneous processing before granting target write authority.

Launch by dependency group, not an arbitrary console selection. Infrastructure services and databases may need to become ready before application workers and schedulers. Validate each gate. Move DNS or Global Accelerator traffic only after application/data readiness, using the cache and connection principles from AWS293.

After traffic moves, watch business transactions, errors, latency, saturation, replication/data state, security, cost, and external integrations. Keep one incident log with actions, evidence, decisions, and deviations.

10. Failback is a second recovery project

Failback must return all target-side changes without creating two writers. Its method depends on destination:

  • reverse replication toward prepared on-premises or other supported infrastructure using the current DRS failback workflow;
  • recovery into a source AWS Region/account pattern supported by current launch settings; or
  • application-native migration when that provides safer data semantics.

Plan network reachability for TCP 443 and TCP 1500 in the reverse direction, target hardware/virtualization, boot compatibility, capacity, agent/service state, time synchronization, credentials, and the failback client's requirements where used. Reverse replication readiness does not prove business consistency.

Run a failback drill. Synchronize until the agreed lag is reached, stop or fence writes in the recovery environment, capture the final consistency point, launch or restore the original-side target, validate it, switch traffic, and preserve a return path. Do not terminate recovery instances or remove recovery data until the retention and acceptance gates pass.

11. Read-only inspection and Linux evidence

Only run these in an authorized account and redact output:

export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text

aws drs describe-source-servers \
  --query 'items[].{Server:sourceServerID,LifeCycle:lifeCycle.state,Replication:dataReplicationInfo.dataReplicationState,Lag:dataReplicationInfo.lagDuration}'

aws drs describe-jobs \
  --query 'items[].{Job:jobID,Type:type,Status:status,InitiatedBy:initiatedBy}'

aws drs describe-recovery-instances \
  --query 'items[].{Recovery:recoveryInstanceID,Source:sourceServerID,State:dataReplicationInfo.dataReplicationState}'

If one source-server ID is explicitly assigned, inspect its replication configuration and launch settings with the matching read operations. Then compare the referenced EC2 launch template, subnet, security groups, IAM profile, and KMS key using read-only service calls.

On a supplied Linux source diagnostic, interpret rather than execute against production:

systemctl status aws-replication-agent
ss -tpn
ip route
getent ahosts drs.ap-south-1.amazonaws.com
nc -zv REPLICATION_SERVER_PRIVATE_IP 1500

A successful TCP connection proves only reachability. It does not prove TLS pairing, throughput, agent health, disk replication, or recoverability. Do not use curl -k, disable firewalls, or open 0.0.0.0/0 casually to hide a certificate, route, or policy failure.

12. Guided three-server recovery workshop

Use the supplied Harbor Claims service:

  • claims-web: Linux/Nginx web tier with uploaded documents on a separate volume;
  • claims-app: Java application with a nightly scheduler and outbound identity/email calls;
  • claims-db: PostgreSQL with 4 TB provisioned disk and average 80 Mbps, peak 220 Mbps changed-block rate;
  • shared target: RTO 90 minutes, RPO 15 minutes;
  • recovery Region: ap-south-1, source outside that Region;
  • available replication path: 300 Mbps effective under normal conditions;
  • known incident: corruption might have started 38 minutes before detection; and
  • drill constraints: no real email, no identity writes, no production DNS, and no duplicate scheduler.

Produce these artifacts:

  1. DRS suitability and rejected-alternative table for each component.
  2. Source/staging/recovery/failback data-flow and trust-boundary diagrams.
  3. Agent prerequisite and source-disk inventory.
  4. Route, firewall, endpoint, proxy, DNS, TCP 443/1500, and return-path matrix.
  5. Average/peak bandwidth, lag-growth, catch-up, initial-sync, and RPO analysis.
  6. IAM, installer, service-role, recovery-instance role, KMS, and emergency-access model.
  7. Replication, DRS launch, EC2 template, and post-launch settings baseline.
  8. Recovery-point investigation timeline with latest-point rejection or acceptance.
  9. Isolated drill design and side-effect controls.
  10. Dependency start/stop order and technical/business validation matrix.
  11. Minute-by-minute recovery runbook and measured RTO ledger.
  12. Source fencing, write-authority, DNS/traffic, and split-brain controls.
  13. Reverse-replication and failback runbook.
  14. Failure injections: port 1500 blocked, S3 endpoint policy denied, lag spike, rescan, wrong subnet, missing IAM profile, failed post-launch action, corrupt latest point, database crash recovery, and quota shortage.
  15. Cost, cleanup, evidence retention, and quarterly game-day plan.

The exercise passes only when all three servers produce one accepted business service. Three running EC2 instances are not enough.

13. Cost, quotas, licenses, and cleanup

Include DRS source-server charges under current terms, replication-server compute, staging EBS volumes sized from provisioned source disks, snapshots and recovery points, KMS, data transfer, NAT or private connectivity, logs, monitoring, drills, recovery EC2/EBS/load balancers, backups, security services, licenses, support, parallel source operation, and failback transfer.

Drills create real resources. Tag them with owner, exercise, retention, and expiration. After evidence is retained, terminate only approved drill instances, remove temporary network/load-balancing/DNS resources, expire temporary credentials and secrets, and confirm source replication remains healthy. Never remove protected sources or uninstall agents merely to reduce a lab bill.

Review quotas for protected source servers, concurrent jobs, replication capacity, snapshots, EC2 instances/vCPUs, EBS volumes/throughput, ENIs, IP addresses, load balancers, KMS requests, and target-service dependencies. Preapprove increases before a disaster.

Check OS and commercial-software licensing in both steady staging and launched recovery. Hardware-bound, host-bound, dedicated-host, or bring-your-own-license rules can change the target design.

Diagnose from layered evidence

SymptomLikely causesNext evidence and correction
Agent not seenService TCP 443, DNS, proxy, credentials, clock, agent serviceVerify process, name resolution, trusted TLS path, installer/service logs
Replication server cannot pairStaging TCP 443/S3/EC2 access or endpoint policyTrace route, DNS, SG/NACL, endpoints and policies
Agent connected but blocks do not flowTCP 1500, route, firewall, certificate pairingTest exact source-to-replication path and inspect DRS error
Lag grows continuouslyThroughput below change rate, throttling, disk bottleneck, rescanCompare change/throughput trend and calculate catch-up
Drill instance does not bootUnsupported layout/driver, wrong point, conversion, launch settingsInspect job, console, volumes, boot mode and serial/system logs
EC2 runs but app failsStartup order, identity, DNS, secrets, certificate, dependency or dataTrace one business transaction through every layer
Latest point reproduces outageCorruption replicatedInvestigate timeline and test an earlier point in isolation
Failback stallsReverse TCP 1500/443, failback client, clock, target capacity or agentValidate reverse path and workflow-specific logs
RTO report looks excellentClock stopped at instance launchMeasure until business acceptance and traffic restoration

Knowledge check

  1. What does DRS replicate?

Changed blocks from supported full servers into staging; it does not understand application transactions.

  1. What are TCP 443 and 1500 used for?

TCP 443 supports agent/staging communication with required AWS services; TCP 1500 carries replication data to staging replication servers.

  1. Why can lag grow even when the agent is connected?

Effective throughput can be lower than the source changed-block rate or a rescan can add work.

  1. What is the difference between a replication server and recovery instance?

The first maintains low-cost staged disk replication; the second is a converted EC2 workload launched for drill or recovery.

  1. Why is the newest recovery point not always best?

It may contain recently replicated corruption, malware, or destructive changes.

  1. When does RTO measurement end?

When the required business service is validated and available, not when an instance reaches running.

  1. Why isolate a recovery drill?

To prevent duplicate writes, messages, schedules, licenses, or customer impact while testing real recovery behavior.

  1. Why does failback need reverse-path testing?

Target-side changes must return through a working, secure replication and consistency process before write authority moves.

Lesson acceptance

You may continue when your submission contains:

  • a defensible DRS fit decision and independent-backup boundary;
  • complete source, agent, TCP 443/1500, staging, snapshot, conversion, and recovery paths;
  • bandwidth, lag, effective-RPO, and business-RTO calculations;
  • least-privilege identities, encryption, emergency access, and audit design;
  • four distinct setting-layer baselines;
  • corruption-aware recovery-point selection;
  • an isolated, measurable three-server drill;
  • one-writer failover and data-safe failback runbooks;
  • application, data, dependency, security, operational, and business evidence;
  • failure-injection diagnoses;
  • full cost, license, quota, retention, and cleanup ownership; and
  • a recurring exercise schedule with tracked improvements.

Official sources

Advertisement