Lesson 189 · AWS Learning Path

AWS 189: Active-active multi-site strategy

· Published · 7 min read

Labelled process diagram for AWS 189: Users in multiple Regions to Global routing to Active regional application and data path to Conflict, isolation, and business-result evidence, with decision, proof and rejection...

Why this lesson matters

Operate more than one site serving traffic while managing data writes, consistency, conflicts, routing, deployment, observability, and partial failure.

What you will be able to do

By the end, you can:

  • explain active-active multi-site strategy in plain language;
  • locate the current service controls in the AWS Management Console;
  • run the matching CloudShell or AWS CLI queries and explain every important field;
  • draw the identity, network, data, failure, and monitoring path;
  • choose the service from requirements and reject it when those requirements are absent;
  • diagnose a failed or misleading result from evidence;
  • state the cost owner and prove cleanup or a no-create result.

Before you start

  • Use a personal AWS account only when its owner has approved the lesson. Do not use the root user for daily work.
  • CloudShell is the default command environment. AWS028 explains CloudShell; AWS029 and AWS030 explain local AWS CLI installation and profiles.
  • The course example Region is ap-south-1. Global services and services with a required control Region are called out in their commands.
  • Run aws sts get-caller-identity privately. Redact the account number before sharing evidence.
  • Never paste access keys, passwords, secret values, private object data, presigned URLs, or full account-specific ARNs into a submission.
  • This is a no-create lesson. Every Console action and AWS CLI command is read-only. Create the practical artifact locally.
  • Console wording can change. Use the Console service search if a menu label has moved, then confirm the current field in the official documentation.

The core model

QuestionWhat it means in this lesson
PurposeOperate more than one site serving traffic while managing data writes, consistency, conflicts, routing, deployment, observability, and partial failure.
Scope and boundaryThe learner must identify the account and Region scope, resource boundary, identity path, data or network path, failure behavior, observability, and cleanup ownership for Active-active multi-site recovery.
Evidence of successSuccess means the Console fields, CLI result, workload behavior, monitoring evidence, and architecture claim agree. An available state alone is not enough for Active-active multi-site recovery.
Cost modelDuplicate regional capacity, global databases, acceleration, transfer, observability, data synchronization, and testing make this the highest-cost DR strategy.
Safe rejection ruleAvoid assuming every database supports safe multi-writer semantics or letting one bad deployment reach all Regions together.

How the request flows

+-----------------------------+
|  Users in multiple Regions  |
+-----------------------------+
              |
              v
+----------------------+
|    Global routing    |
+----------------------+
           |
           v
+---------------------------------------------+
|  Active regional application and data path  |
+---------------------------------------------+
                      |
                      v
+-----------------------------------------------------+
|  Conflict, isolation, and business-result evidence  |
+-----------------------------------------------------+

Active-active changes the application design

Multi-site active-active serves production traffic from multiple Regions simultaneously. It can approach low disruption for a regional infrastructure failure, but it does not make every operation strongly consistent or eliminate recovery. Corruption, global dependencies, bad deployments and stolen credentials can affect every Region; backups and rollback remain mandatory.

Partition users and writes deliberately

Data patternBenefitRisk/design requirement
Region-local ownershipSimple writes within a home RegionRoute each entity consistently and define ownership transfer during failure.
Single global writerEasier consistencyWriter Region remains a dependency and promotion has nonzero RTO/RPO.
Multi-writer eventual consistencyLocal writes and availabilityDefine conflict resolution, idempotency, ordering and user-visible convergence.
Synchronous cross-Region coordinationStronger consistencyWAN latency and partition availability can make it unsuitable for request path.

DynamoDB global tables use multi-Region replication with conflict behavior that the application must accept. Aurora Global Database normally has one writer Region with cross-Region readers and managed switchover/failover patterns; it is not automatically multi-writer. S3 replication is asynchronous and version/delete-marker behavior must be configured. Choose from the transaction semantics, not the “global” label.

Globally unique idempotency keys, conditional writes, monotonic/version metadata, event deduplication and reconciliation jobs are application responsibilities. Test concurrent updates and network partitions, not only Region shutdown.

Traffic and capacity

Route 53 latency/geolocation/weighted/failover policies or Global Accelerator direct users among healthy regional endpoints. Health must represent business readiness and data-write eligibility. Removing one Region requires every survivor to absorb its peak traffic, connections, queues and database load; N+1 tested capacity and quotas are essential.

Sessions should be stateless or stored in an intentionally replicated system. Region-local cookies, caches, WebSocket connections and in-flight jobs will not teleport. Clients need bounded retries and idempotency so routing changes do not duplicate orders.

Deployment and blast radius

Do not deploy a bad release to all Regions at once. Use waves/canaries, independent regional rollback and feature flags with a safe default. Keep DNS, identity, CI/CD, artifact, KMS, observability and secrets dependencies from becoming hidden single-Region failures. Separate credentials/accounts where compromise isolation requires it.

Active-active DR tests remove one Region from traffic and prove surviving capacity, data behavior, background processing and recovery. Data-disaster tests restore a clean point and reconcile separately because infrastructure availability cannot undo corruption.

Architecture decision table

SituationDirectionReason
Requirement matchesUse active-active only when recovery and global service requirements justify its data and operational complexity.Select only after scope, behavior, security, recovery, operations, and price evidence agree.
Requirement does not matchAvoid assuming every database supports safe multi-writer semantics or letting one bad deployment reach all Regions together.Rejecting an attractive service is a valid architecture result.
No create permission or cost approvalUse supplied evidence and local design workLearning does not depend on creating an hourly resource.
Existing resource is unknown or unownedInspect only, then stopNever change or delete a resource merely because it resembles a course example.

AWS Management Console, step by step

Sign in with the normal non-root learning identity. Write the expected starting state before opening the service.

  1. Use the Console service search and open Route 53 or Global Accelerator, CloudFront, Global Tables or Aurora Global Database, and regional application inventories; confirm the account and Region before reading the page.
  2. Inspect the supplied or owned resource's status, configuration, permissions, networking, encryption, monitoring, tags, and dependencies without changing it.
  3. Open the related metrics, logs, events, or history view and record one timestamped signal that would prove or disprove the expected behavior.
  4. Return to the resource list, clear filters, and record the final inventory. On the read-only track, do not choose Create, Save, or Delete.

CloudShell and AWS CLI, step by step

Start with a known caller and Region:

export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
aws configure list

Redact the account part of the ARN in shared evidence. Now run the topic queries:

aws route53 list-health-checks --output table
aws globalaccelerator list-accelerators --region us-west-2 --output table
aws dynamodb list-global-tables --output table
aws rds describe-global-clusters --output table

Expected interpretation

Two running Regions are not active-active unless the application serves valid traffic and data behavior is safe from both. Partial dependency and stale-write failure must be tested.

Practical work

Design active-active order status, then decide whether order creation can also be active-active. Specify write ownership, idempotency, conflict handling, session state, global routing, deployment waves, observability, isolation, and failback.

Diagnose this topic from its own evidence

  • Users see stale/conflicting writes: identify writer ownership, replication lag, conflict rule and client routing.
  • Region removed but global errors rise: inspect survivor quotas/capacity and global shared dependencies.
  • Duplicate transactions after retry: trace idempotency key, queue semantics and cross-Region deduplication store.
  • Healthy endpoint should not write: separate liveness from write-readiness and data-role health.
  • Bad release affects all Regions: inspect deployment wave/rollback isolation and globally shared configuration.

Cost and cleanup

Duplicate regional capacity, global databases, acceleration, transfer, observability, data synchronization, and testing make this the highest-cost DR strategy.

Knowledge check

  1. What operational purpose is this lesson solving?

Expected direction: Operate more than one site serving traffic while managing data writes, consistency, conflicts, routing, deployment, observability, and partial failure.

  1. Which scope or ownership boundary must be proved first?

Expected direction: The learner must identify the account and Region scope, resource boundary, identity path, data or network path, failure behavior, observability, and cleanup ownership for Active-active multi-site recovery.

  1. What evidence is strong enough to accept the result?

Expected direction: Success means the Console fields, CLI result, workload behavior, monitoring evidence, and architecture claim agree. An available state alone is not enough for Active-active multi-site recovery.

  1. Which tempting design or shortcut must be rejected?

Expected direction: Avoid assuming every database supports safe multi-writer semantics or letting one bad deployment reach all Regions together.

  1. Which cost dimensions and retained resources need an owner?

Expected direction: Duplicate regional capacity, global databases, acceleration, transfer, observability, data synchronization, and testing make this the highest-cost DR strategy.

Lesson acceptance

  • Choose a data ownership/consistency model and explain partition behavior.
  • Prove survivors can absorb full failure traffic and dependencies.
  • Design health-aware global routing without assuming sessions/connections move.
  • Demonstrate idempotent retry, conflict resolution and reconciliation.
  • Include staggered deployment, corruption recovery, backup and regional re-entry tests.

Official sources

Advertisement