AWS 190: Multi-AZ and Multi-Region architecture
Why this lesson matters
Use AZs for high availability inside one Region and Regions for disaster recovery, sovereignty, or global needs, with different latency and data assumptions.
What you will be able to do
By the end, you can:
- explain multi-az and multi-region architecture in plain language;
- locate the current service controls in the AWS Management Console;
- run the matching CloudShell or AWS CLI queries and explain every important field;
- draw the identity, network, data, failure, and monitoring path;
- choose the service from requirements and reject it when those requirements are absent;
- diagnose a failed or misleading result from evidence;
- state the cost owner and prove cleanup or a no-create result.
Before you start
- Use a personal AWS account only when its owner has approved the lesson. Do not use the root user for daily work.
- CloudShell is the default command environment. AWS028 explains CloudShell; AWS029 and AWS030 explain local AWS CLI installation and profiles.
- The course example Region is
ap-south-1. Global services and services with a required control Region are called out in their commands. - Run
aws sts get-caller-identityprivately. Redact the account number before sharing evidence. - Never paste access keys, passwords, secret values, private object data, presigned URLs, or full account-specific ARNs into a submission.
- This is a no-create lesson. Every Console action and AWS CLI command is read-only. Create the practical artifact locally.
- Console wording can change. Use the Console service search if a menu label has moved, then confirm the current field in the official documentation.
The core model
| Question | What it means in this lesson |
|---|---|
| Purpose | Use AZs for high availability inside one Region and Regions for disaster recovery, sovereignty, or global needs, with different latency and data assumptions. |
| Scope and boundary | The learner must identify the account and Region scope, resource boundary, identity path, data or network path, failure behavior, observability, and cleanup ownership for Multi-AZ and Multi-Region architecture. |
| Evidence of success | Success means the Console fields, CLI result, workload behavior, monitoring evidence, and architecture claim agree. An available state alone is not enough for Multi-AZ and Multi-Region architecture. |
| Cost model | Requests, running capacity, storage, logs, data transfer, retained state, and optional features must be priced for the exact design. |
| Safe rejection rule | Avoid presenting Multi-AZ as regional disaster recovery or assuming service configuration automatically replicates every dependency cross-Region. |
How the request flows
+-----------------------------+
| Local or regional failure |
+-----------------------------+
|
v
+------------------------------+
| Multi-AZ regional recovery |
+------------------------------+
|
v
+-----------------------------------------------+
| Regional disaster and cross-Region recovery |
+-----------------------------------------------+
|
v
+-----------------------------+
| Business RPO/RTO evidence |
+-----------------------------+
Failure scope determines architecture
An Availability Zone is one or more discrete data centers with independent infrastructure inside a Region. Multi-AZ design targets an AZ/component failure with regional services and low-latency networking. Multi-Region design targets regional disruption, geography, sovereignty or global latency, adding independent control planes, asynchronous data decisions, cost and operational complexity.
| Requirement/failure | Minimum direction | Why |
|---|---|---|
| Instance/host failure | Auto Scaling or service replacement | More AZs do not help if no replacement mechanism exists. |
| AZ loss | Subnets/targets/capacity across AZs and an AZ-resilient data tier | A load balancer cannot help if all healthy capacity/data is in one AZ. |
| Regional control/data-plane loss | Tested recovery or active workload in another Region | Multi-AZ remains inside the affected Region. |
| Data corruption/deletion | Versioned immutable backups and restore | Replication can copy corruption across AZs/Regions. |
| Low global user latency | Multi-Region serving/edge architecture | DR standby alone may not serve users normally. |
| Regulatory location/isolation | Explicit account/Region/data controls | Availability patterns do not prove compliance. |
Build a Multi-AZ workload correctly
Use at least two AZ subnets for ALB/compute, distribute desired/minimum capacity, rebalance after failures and design each surviving AZ for required load. NAT gateways and interface endpoints are zonal cost/availability decisions; avoid cross-AZ dependencies accidentally. RDS Multi-AZ provides standby/failover according to deployment type, not read scaling by default. EFS is regional with zonal mount targets; EBS volumes are AZ-scoped and snapshots are regional. Map each service boundary explicitly.
Fault isolation includes application behavior: stateless compute, external session state, bounded timeouts, retries with jitter, idempotency, backpressure and health checks that remove failed capacity. Cross-AZ charges and quorum behavior belong in the design.
Add a Region only for a justified objective
Select recovery Region from service/feature availability, latency, data residency, quotas, capacity, inter-Region dependencies and organizational policy - not nearest name. Replicate IaC, artifacts, configuration, certificates, secrets and keys deliberately. Most resource ARNs and service states are regional; copying data alone does not create a runnable application.
Classify each dependency as regional, global, replicated or external. Route 53 and IAM have global aspects but still have data/control-plane behaviors; do not assume “global service” means no failure mode. Minimize recovery-time control-plane calls for strict RTO and pre-create critical resources.
Prove resilience with failure injection and game days
Run component and AZ tests before regional exercises. Define steady state, hypothesis, safety boundaries, abort conditions and rollback. Measure user-visible availability, queue depth, data loss/lag, capacity recovery and alarms. Regional DR evidence includes declaration through business validation and traffic movement, not a screenshot of a replica.
Architecture decision table
| Situation | Direction | Reason |
|---|---|---|
| Requirement matches | Use Multi-AZ as the normal regional availability foundation and add Multi-Region only for explicit business requirements. | Select only after scope, behavior, security, recovery, operations, and price evidence agree. |
| Requirement does not match | Avoid presenting Multi-AZ as regional disaster recovery or assuming service configuration automatically replicates every dependency cross-Region. | Rejecting an attractive service is a valid architecture result. |
| No create permission or cost approval | Use supplied evidence and local design work | Learning does not depend on creating an hourly resource. |
| Existing resource is unknown or unowned | Inspect only, then stop | Never change or delete a resource merely because it resembles a course example. |
AWS Management Console, step by step
Sign in with the normal non-root learning identity. Write the expected starting state before opening the service.
- Use the Console service search and open EC2 Availability Zones, RDS placement, load balancers, and cross-Region resources; confirm the account and Region before reading the page.
- Inspect the supplied or owned resource's status, configuration, permissions, networking, encryption, monitoring, tags, and dependencies without changing it.
- Open the related metrics, logs, events, or history view and record one timestamped signal that would prove or disprove the expected behavior.
- Return to the resource list, clear filters, and record the final inventory. On the read-only track, do not choose Create, Save, or Delete.
CloudShell and AWS CLI, step by step
Start with a known caller and Region:
export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
aws configure list
Redact the account part of the ARN in shared evidence. Now run the topic queries:
aws ec2 describe-availability-zones --query 'AvailabilityZones[?ZoneType==`availability-zone`].{Name:ZoneName,Id:ZoneId,State:State}' --output table
aws rds describe-db-instances --query 'DBInstances[].{Id:DBInstanceIdentifier,AZ:AvailabilityZone,MultiAZ:MultiAZ}' --output table
aws ec2 describe-regions --all-regions --query 'Regions[].{Region:RegionName,Status:OptInStatus}' --output table
Expected interpretation
AZ design addresses local failures with low-latency regional services. Multi-Region adds separate control planes, data replication, routing, quotas, deployments, costs, and failure modes.
Practical work
Draw a two-AZ primary Region and one recovery Region. Place load balancers, app capacity, database, backups, replication, keys, DNS, CI/CD, monitoring, and operators. Mark synchronous and asynchronous paths.
Diagnose this topic from its own evidence
- Multi-AZ but outage persists: find single-AZ database, NAT, endpoint, capacity, state or deployment dependency.
- Cross-AZ traffic/cost spikes: inspect subnet placement, load-balancer cross-zone behavior and zonal service endpoints.
- Recovery Region cannot launch: inspect quotas, service availability, artifacts, keys, certificates, IAM and IaC parameters.
- Failover creates stale data: compare replication checkpoint/lag with RPO and writer fencing.
- Test passes infrastructure but users fail: include identity, DNS/TLS, dependencies and business transaction validation.
Cost and cleanup
Requests, running capacity, storage, logs, data transfer, retained state, and optional features must be priced for the exact design.
Knowledge check
- What operational purpose is this lesson solving?
Expected direction: Use AZs for high availability inside one Region and Regions for disaster recovery, sovereignty, or global needs, with different latency and data assumptions.
- Which scope or ownership boundary must be proved first?
Expected direction: The learner must identify the account and Region scope, resource boundary, identity path, data or network path, failure behavior, observability, and cleanup ownership for Multi-AZ and Multi-Region architecture.
- What evidence is strong enough to accept the result?
Expected direction: Success means the Console fields, CLI result, workload behavior, monitoring evidence, and architecture claim agree. An available state alone is not enough for Multi-AZ and Multi-Region architecture.
- Which tempting design or shortcut must be rejected?
Expected direction: Avoid presenting Multi-AZ as regional disaster recovery or assuming service configuration automatically replicates every dependency cross-Region.
- Which cost dimensions and retained resources need an owner?
Expected direction: Requests, running capacity, storage, logs, data transfer, retained state, and optional features must be priced for the exact design.
Lesson acceptance
- Map every component's AZ/Region/global scope and failure behavior.
- Design survivor capacity, zonal networking and state for an AZ loss.
- Justify Multi-Region complexity with a business RTO/RPO, latency or regulatory need.
- Inventory and pre-stage all non-data recovery dependencies.
- Produce measured game-day evidence with steady state, abort, rollback and business validation.