AWS 035: Regions, Availability Zones and edge locations
The problem
A workload is deployed across two Availability Zones and described as disaster recovery. Static content is cached at edge locations and the team assumes the database is now global. The location terms are correct, but the conclusions are not.
This lesson turns location definitions into architecture decisions.
Final outcome
You will place a three-tier workload across Regions, Availability Zones, and edge locations, identify which failures each layer addresses, and trace normal and failure traffic without creating AWS resources.
Three distinct purposes
| Layer | Primary architecture purpose | Does not automatically provide |
|---|---|---|
| Region | geographic, regulatory, and broad isolation boundary | cross-Region recovery or replication |
| Availability Zone | fault isolation within a Region | protection from regional disruption |
| Edge location | low-latency delivery and supported edge processing | arbitrary VPC compute or database replication |
Using all three does not guarantee availability. Dependencies and state determine the outcome.
Reference application
The training portal has:
- cacheable course images and documents;
- dynamic learner requests;
- assessment writes;
- identity dependency;
- transactional data;
- logs and audit records;
- RTO of two hours;
- RPO of 15 minutes.
Regional placement
Choose one primary Region because it satisfies data location, service availability, connectivity, operational skills, and cost. Select a recovery Region only if the recovery strategy requires it.
Document which data can cross the regional boundary. Replication must address:
- encryption and keys;
- identity and secrets;
- conflict or failover behavior;
- data transfer;
- lag and RPO;
- capacity in recovery;
- DNS or traffic switching;
- testing and failback.
Multi-AZ placement
Within the primary Region:
regional entry
/ \
v v
AZ A app AZ B app
\ /
v v
resilient data
Avoid placing every application instance and its only database copy in one zone. A zonal design also needs:
- independent subnets and route paths;
- capacity in each remaining zone;
- health checks that represent real service;
- state outside replaceable application workers;
- deployment that does not break all zones together;
- data service configured for the required resilience;
- tested client retry and reconnection.
If each zone depends on one self-managed appliance in AZ A, the diagram is still effectively single-zone.
Edge placement
Place cacheable static content behind a CDN when requirements justify it. The edge can:
- cache objects near users;
- terminate supported TLS and HTTP behavior;
- reduce repeated origin requests;
- apply supported security and request logic;
- route to an origin according to service configuration.
The first request or an expired object may still go to the origin. Dynamic or personalized content needs deliberate caching rules. Never cache private learner data under a shared cache key.
Trace three paths
Cache hit
learner -> nearby edge -> cached approved object -> learner
The primary application may not receive the request.
Cache miss
learner -> edge -> regional origin -> object response
<- edge stores according to policy <-
Origin availability, authorization, and cache-control policy matter.
Assessment submission
learner -> entry -> healthy application in one AZ
-> transactional data path -> durable acknowledgement
Do not route state-changing assessment writes as if they were static cache objects.
Failure analysis
| Failure | Expected design response |
|---|---|
| one app process fails | health check removes it and replacement starts |
| one AZ fails | capacity and state in another AZ continue |
| primary Region fails | tested recovery strategy activates if required |
| edge cache entry expires | request returns to healthy origin |
| origin unavailable | only suitable cached objects may remain available |
| replication lags | recovery point may lose recent data within measured RPO |
| bad deployment reaches both AZs | zonal redundancy may not help |
Availability Zones reduce correlated infrastructure failure. They do not protect from every shared software, identity, configuration, or data error.
Practical architecture artifact
Create:
mkdir -p "$HOME/nitwings-aws/evidence/aws-035"
Create location-architecture.md with:
- primary and recovery Region requirements;
- two or more Availability Zones in the primary Region;
- public entry, private application, and data placement;
- edge-cached and noncached request classifications;
- state and replication paths;
- normal request traces;
- AZ failure trace;
- Region recovery trace;
- RTO/RPO evidence points;
- cost and data-transfer consequences.
Use labels such as AZ ID <zone-a> rather than assuming one account's ap-south-1a maps physically to another account's same letter.
Inspect the account map
Run:
COURSE_REGION="ap-south-1"
aws ec2 describe-availability-zones \
--profile course \
--region "$COURSE_REGION" \
--query 'AvailabilityZones[?ZoneType==`availability-zone`].{Name:ZoneName,Id:ZoneId,State:State}' \
--output table \
--no-cli-pager
Copy only zone names and IDs into the private artifact. This query proves the account's visible ordinary Availability Zones and state. It does not prove capacity for a chosen instance type or service feature.
For edge locations, use current AWS global infrastructure and CloudFront documentation. There is no general EC2 command that turns edge points of presence into normal Availability Zones.
Latency exercise
Do not use one laptop ping as the only Region decision. A useful latency study includes:
- representative user geographies;
- DNS and TLS setup;
- application request latency;
- internet-provider variation;
- origin and dependency location;
- repeated samples;
- percentile distribution;
- failure and congestion cases.
CloudFront can reduce delivery latency for cacheable content, while write latency still depends on the regional application and data path.
Certification decision patterns
- Need protection from one data-center or zonal failure: use multi-AZ architecture.
- Need geographic disaster recovery or data locality: evaluate multi-Region.
- Need low-latency global delivery of cacheable content: evaluate CloudFront edge caching.
- Need compute near a metro area with supported services: evaluate Local Zones.
- Need telecommunications 5G edge placement: evaluate Wavelength.
Select only the mechanism that satisfies the stated requirement. More locations add cost and operational complexity.
Common mistakes
- calling Multi-AZ a backup;
- treating an edge cache as authoritative data;
- using identical AZ letters as cross-account alignment;
- choosing recovery Region without capacity or keys;
- assuming every service supports the same replication;
- ignoring inter-AZ and inter-Region transfer;
- caching authenticated responses without correct cache keys and policy;
- testing failover without testing failback.
Knowledge check
- Does Multi-AZ protect automatically from a bad application release?
- Is an edge location an ordinary subnet location?
- What does RPO measure during Region recovery?
- Why use AZ IDs across accounts?
- Can a cache hit avoid the origin?
Expected answers: no; no; acceptable data loss point; physical alignment; yes for a valid cached object.
Completion gate
Pass when the architecture correctly separates Region, AZ, and edge purposes; traces cache hit, miss, write, AZ failure, and Region recovery; aligns zones with IDs; and links every placement to requirements, evidence, and cost.
No resources were created.