AWS 275: Multi-Region network architecture
Why this lesson matters
Deploying a second copy of a VPC does not create a recoverable multi-Region application. Clients need a healthy entry point, networks need unique routes, services need Regional dependencies, data needs a consistency and conflict model, identities and secrets must be available, and operators need a recovery process that still works when a control plane is impaired.
Multi-Region architecture buys a larger fault-isolation boundary only when the workload can run independently in another Region. It also introduces inter-Region transfer, replication lag, static routes or global policy, certificate/DNS coordination, security duplication, operational drift, and difficult failure modes. This lesson designs the network as one part of an application recovery system rather than treating global routing as disaster recovery.
Outcomes
By the end, you can:
- derive network requirements from business impact, recovery time objective (RTO), and recovery point objective (RPO);
- distinguish backup/restore, pilot light, warm standby, active/passive, and active/active;
- create a non-overlapping Regional IPv4/IPv6 address and ASN plan with IPAM;
- compare TGW inter-Region peering, AWS Cloud WAN, private endpoints, and application-layer patterns;
- trace Regional, inter-Region, hybrid, DNS, ingress, egress, inspection, and return paths;
- choose Route 53, Global Accelerator, CloudFront, or application discovery by protocol and failure behavior;
- design Regional DNS, Resolver endpoints/rules, private hosted zones, and service discovery;
- separate data-plane continuity from control-plane automation;
- model data residency, encryption, identity, secrets, logging, and security policy;
- test Region isolation, partial failure, stale health, replication lag, and failback; and
- calculate inter-Region network, replication, endpoint, inspection, and standby cost.
Begin with the business outcome
Ask what business operation must continue, for whom, at what reduced capacity, after which failures, and with how much data loss. Then define:
| Term | Question |
|---|---|
| RTO | How long may the business capability remain unavailable? |
| RPO | How much committed data may be lost or require reconciliation? |
| Recovery capacity | What percentage of normal traffic must the alternate Region serve? |
| Isolation scope | AZ, Region, dependency, operator error, identity, or data corruption? |
| Consistency | Can clients read stale data or create conflicting writes? |
| Residency | In which jurisdictions may data, logs, backups, keys, and support access exist? |
| Failover authority | Who or what declares disaster and shifts traffic? |
| Failback | How are data and traffic safely returned after recovery? |
A 60-second DNS or accelerator failover cannot meet a 60-second RTO if the database needs two hours to restore. A zero-RPO requirement cannot be claimed with asynchronous replication that has measurable lag. The network design must use the application's realistic readiness.
Recovery patterns
| Pattern | Alternate Region state | Typical tradeoff |
|---|---|---|
| Backup and restore | infrastructure/data restored after event | lowest standing cost, longest RTO/RPO |
| Pilot light | critical data/core services running | faster recovery, substantial scale/configuration work remains |
| Warm standby | complete reduced-capacity stack | shorter RTO, ongoing capacity and drift management |
| Active/passive | secondary ready but not serving normal writes | simpler data ownership, idle/standby cost |
| Active/active reads | reads served in many Regions, writes controlled | latency gains with clearer write authority |
| Active/active writes | several Regions accept writes | fastest local writes, hardest conflict/consistency model |
Active/active is not automatically more resilient. A bad deployment, corrupted replicated data, global credential, DNS policy, or shared dependency can fail every Region at once. Sometimes isolated active/passive plus tested promotion provides better safety.
Classify every dependency as Regional, global data plane, global control plane, external, or organizational. Record whether it is replicated, cached, reconstructed, replaced, or intentionally unavailable during recovery.
Addressing and IPAM
Allocate unique summarizable CIDRs per Region, environment, and trust zone. For example:
| Region | Aggregate | Production | Nonproduction | Shared |
|---|---|---|---|---|
| Region A | 10.64.0.0/12 | 10.64.0.0/14 | 10.68.0.0/14 | 10.72.0.0/16 |
| Region B | 10.80.0.0/12 | 10.80.0.0/14 | 10.84.0.0/14 | 10.88.0.0/16 |
| Region C | 10.96.0.0/12 | 10.96.0.0/14 | 10.100.0.0/14 | 10.104.0.0/16 |
These are planning examples, not allocations to copy blindly. Check corporate, partner, acquisition, VPN, Direct Connect, container/pod, service, link-local, multicast, and IPv6 ranges. Summaries must not include prefixes routed somewhere else.
VPC IP Address Manager (IPAM) uses scopes, pools, and allocations. Build top-level private pools, Regional child pools with the correct locale, then environment/application pools. An IPAM pool allocates only in its locale. Select operating Regions deliberately, share pools with AWS RAM, enforce allocation tags/netmask ranges, and monitor overlap/compliance. Separate scopes can reuse addresses only when networks are guaranteed never to connect; that assumption often expires after mergers or analytics integration.
Plan IPv6 globally as well as IPv4. Separate IPv6 routing, egress, firewall, load balancer, DNS AAAA, client preference, and failure behavior. An alternate Region that works only over IPv4 is not dual-stack recovery.
Use unique TGW autonomous system numbers where possible. Although TGW peering uses static routes today, unique ASNs avoid ambiguity and support future/hybrid evolution.
Inter-Region connectivity choices
Transit Gateway peering
Peer Regional TGWs over the AWS global network. Peering supports static routes only: routes do not propagate across the peering attachment, and peering does not provide ECMP. Each relevant source-associated TGW route table needs a static route to the remote Regional summary, and the peer needs return routes. VPC subnet route tables still need their TGW routes.
Advantages include explicit segmentation and straightforward Regional hubs. Risks include route growth, manual/static lifecycle, hidden one-sided routes, summary mistakes, and operational scaling. TGW peering also does not make Route 53 Resolver in one Region transparently resolve names across the peer; design DNS endpoints/rules per Region.
AWS Cloud WAN
Cloud WAN provides a policy-managed global core network with segments, attachment policies, and Regional core network edges. It is useful when many Regions, sites, and SD-WAN attachments need consistent global segmentation and centralized policy. It does not eliminate VPC route tables, application recovery, DNS, data replication, or policy testing. A core network policy error can have broad blast radius, so version, review, simulate, and roll out policy changes.
Application and service-level connectivity
Not every workload needs routed inter-Region VPC connectivity. Public/edge APIs with TLS and identity, PrivateLink cross-Region endpoint services, asynchronous queues/events, database replication, object replication, or VPC Lattice patterns can expose only required services. Narrow service connectivity reduces route blast radius and overlap concerns.
| Requirement | Starting choice |
|---|---|
| Few Regional hubs and explicit routes | TGW peering |
| Many Regions/sites with global segmentation policy | Cloud WAN |
| One private provider service | PrivateLink/cross-Region service pattern |
| Loose application integration | queue/event/API/object replication |
| No inter-Region runtime dependency | independent Regional stacks |
Avoid circular runtime dependencies where Region A's application requires Region B's identity, DNS, firewall, logging, or database and Region B requires A. During partition, both can fail.
Trace three independent network planes
Client ingress
Internet or corporate clients must reach the selected healthy Region. Trace public DNS/accelerator, edge network, WAF/firewall, Regional load balancer/API, application, and return. Preserve client IP requirements, TLS name/certificate, IPv4/IPv6, sticky sessions, and long-lived connection behavior.
Regional east-west and egress
Each Region should have local DNS, inspection, NAT/egress, endpoints, and dependencies needed during isolation. Sending Region B's internet egress or logs through Region A creates a hidden Regional dependency. If policy deliberately centralizes something, state the failure mode and reduced service.
Inter-Region/hybrid
Trace every TGW/Cloud WAN route, attachment, inspection point, DX/VPN path, customer route, and reverse lookup. On-premises can advertise/receive multiple Regional prefixes, but BGP preference and asymmetric return paths need proof. A Region failover may overload VPN backup or customer firewalls sized only for normal traffic.
For every critical flow write both directions, including DNS and control calls. Use the AWS271 route-proof method and AWS274 inspection method.
Route 53 versus Global Accelerator
Route 53 returns DNS answers based on routing policy and health evaluation. Policies include failover, latency, geolocation, geoproximity, weighted, and multivalue behavior. Clients and intermediate resolvers cache answers for TTL; existing connections do not move merely because DNS changes. Health checks must measure the business-critical path, not only a load balancer listener.
Global Accelerator provides static anycast IP addresses and sends TCP/UDP traffic over the AWS global network to Regional endpoints. A standard accelerator chooses based on client location, endpoint health, traffic dials, and endpoint weights. New connections can move quickly when health changes; existing TCP sessions can still fail and reconnect. If no endpoints are healthy, Global Accelerator can route to all endpoints, so health and application failure behavior must be understood.
| Need | Route 53 | Global Accelerator |
|---|---|---|
| DNS-based web or service selection | strong fit | can complement |
| static global IP allowlisting | not provided by DNS alone | strong fit |
| TCP/UDP non-HTTP acceleration | DNS only selects address | strong fit |
| fine geographic DNS policy | many policy types | proximity/health endpoint groups |
| shift new traffic without waiting full DNS cache | constrained by resolver/client TTL | traffic dial/health acts at accelerator |
| existing connection migration | no | no transparent session migration |
CloudFront is often the global entry for cacheable/content and HTTP workloads, with origin failover and WAF. API Gateway, ALB, and application-layer routers may fit other cases. Select based on protocol, static-IP need, caching, WAF, origin behavior, source preservation, TLS, and cost.
Global Accelerator traffic dial applies to traffic already directed to an endpoint group, not a percentage of all global traffic. Endpoint weights divide traffic within a group. Source-IP affinity can improve stickiness but can create uneven distribution and does not solve cross-Region session/data state.
Health, traffic shifting, and data readiness
A health signal must represent whether the Region can complete a safe transaction. Combine infrastructure and synthetic application checks carefully. A health check that writes production data may cause side effects; a shallow /health may stay green while database, identity, DNS, or queue dependency is broken.
Use readiness gates before traffic shift:
- infrastructure and capacity deployed;
- latest approved application/configuration version;
- secrets, certificates, keys, and feature flags available;
- data replication lag within RPO and promotion status known;
- local DNS, endpoints, egress, inspection, logging, and identity working;
- quotas and scaling tested for recovery load;
- synthetic read and controlled write/rollback pass; and
- incident commander has explicit go/no-go evidence.
Use weighted DNS or accelerator dials for a canary, observe errors/latency/replication, then increase gradually. Automatic failover is appropriate only when signals are reliable and failover cannot cause split-brain writes. Otherwise automate evidence collection and require controlled approval.
Data, state, and consistency
Network reachability cannot make a secondary database writable safely. For each store record replication type, direction, lag, consistency, conflict resolution, failover/promotion, endpoint discovery, backup independence, encryption key, and failback/resynchronization.
Read-local/write-primary patterns reduce write conflicts but make the primary Region a dependency. Global databases can support multi-Region reads or writes with service-specific semantics. S3 cross-Region replication is asynchronous and existing-object replication/versioning/deletion behavior must be designed. Queues, caches, search indexes, file systems, and object stores all have different replication guarantees.
Sessions should be portable, reconstructible, or deliberately reauthenticated. Do not rely on load-balancer stickiness to preserve a session after Regional failover. Idempotency keys, globally unique identifiers, clocks, duplicate event handling, and conflict rules belong in the recovery design.
Data residency includes replicas, backups, snapshots, logs, traces, DNS/query records, security findings, and support access. Encryption in transit and at rest does not by itself satisfy location restrictions.
DNS and service discovery
Create Regional Route 53 Resolver inbound/outbound endpoints and rules when each Region must operate independently. Share/associate private hosted zones and Profiles according to Region/account behavior. Test split-view DNS, negative caching, stale records, and private-zone conflicts.
Use Regional service names such as api.region-a.internal.example plus a controlled global name when helpful. A global private name should not direct clients to a Region whose data is not ready. Service discovery records need health and lifecycle semantics; DNS registration alone is not application readiness.
TGW peering does not provide cross-Region Amazon-provided DNS resolution. Forwarding DNS across a peering link to one central Region recreates a dependency. Prefer local Resolver paths with deliberate cross-Region rules only where required.
Inspection and security independence
Deploy Regional inspection, NAT, DNS filtering, VPC endpoints, WAF, logging buffers, and security controls needed for standalone operation. Replicate policies through versioned automation, but permit controlled Regional rollout to limit global blast radius.
Central policy services such as Firewall Manager or Cloud WAN simplify governance but their data-plane and control-plane failure behavior must be documented. Existing Regional resources may continue processing while configuration APIs are unavailable. Do not make failover require creating a new firewall, endpoint, certificate, or IAM role during an incident if the RTO cannot tolerate that dependency.
Use separate break-glass roles, tested access paths, hardware MFA/process controls, and audit destinations. Replicate secrets only to approved Regions using service-supported mechanisms. KMS keys are Regional resources even when related multi-Region key features are used; policies, grants, aliases, and service integration still require proof.
Control plane versus data plane
The data plane processes established service traffic. The control plane creates or changes DNS records, routes, accelerators, scaling, certificates, and policies. A resilient design pre-provisions its recovery data plane so routine failover mainly changes a small, tested steering control.
Avoid requiring a long chain of control-plane calls during failure. Examples include creating a VPC, requesting quota, restoring every secret, provisioning endpoints, issuing certificates, and then changing DNS. Pre-create and continuously validate what the RTO needs.
Route 53 is a global service whose control plane has Regional hosting details; its globally distributed authoritative data plane is separate. Do not equate one console/API impairment with DNS data-plane failure. Cache emergency runbooks and approved infrastructure definitions outside the affected operational path.
Read-only evidence collection
Inventory every enabled Region explicitly:
aws sts get-caller-identity --query Arn --output text
aws ec2 describe-regions + --query 'Regions[].{Region:RegionName,OptIn:OptInStatus}' --output table
for region in ap-south-1 ap-southeast-1 eu-west-1; do
aws ec2 describe-transit-gateways --region "$region" --output json
aws ec2 describe-transit-gateway-peering-attachments + --region "$region" --output json
aws ec2 describe-vpcs --region "$region" + --query 'Vpcs[].{Id:VpcId,Cidrs:CidrBlockAssociationSet[].CidrBlock}' + --output json
done
Global and multi-Region services:
aws route53 list-health-checks --output json
aws globalaccelerator list-accelerators --region us-west-2 --output json
aws networkmanager list-core-networks --region us-west-2 --output json
aws ec2 describe-ipams --region ap-south-1 --output json
aws ec2 describe-ipam-pools --region ap-south-1 --output json
Global Accelerator API calls use us-west-2. Network Manager/Cloud WAN command Regions and resource scope must be confirmed for the deployed design. Also inspect TGW routes/associations/propagations, VPC route tables, Cloud WAN policy versions/routes, Route 53 records/health status, Resolver, WAF/firewalls, load balancers, DX/VPN/BGP, quotas, replication metrics, CloudTrail, Config, CUR, and application SLOs. Redact identities, IP plans, domains, routes, and recovery topology.
Failure diagnosis
| Symptom | First proof | Common cause |
|---|---|---|
| DNS selects unhealthy Region | health-check chain and TTL | shallow/stale health or cached answer |
| Accelerator stays in Region | endpoint/endpoint-group health | healthy LB but broken dependency |
| New Region reachable, writes fail | database promotion/lag | network shifted before data readiness |
| One-way inter-Region flow | static TGW routes both sides | missing return route |
| Remote hostname fails | Regional Resolver design | assumed DNS across TGW peering |
| On-premises uses old Region | BGP advertisements/preferences | route not withdrawn or customer preference |
| Recovery overloads | quotas/capacity metrics | standby never load-tested |
| Only IPv6 clients fail | AAAA/IPv6 route/security | partial dual-stack deployment |
| Security logs disappear | Regional destination/dependency | logs depended on failed Region |
| Failback duplicates data | replication/conflict evidence | no reconciliation/idempotency process |
Start at client resolution/accelerator selection, then Regional ingress, application dependencies, data authority, inter-Region and hybrid routes, security path, and return. Correlate UTC evidence from both Regions. Do not change DNS and routes simultaneously without knowing which change corrected the fault.
Game days, failover, and failback
Test isolated dependency failure before whole-Region simulation. Include load balancer health, DNS Resolver, firewall/NAT, TGW/Cloud WAN route, DX, identity, secrets, database lag, logging destination, control-plane API, and operator access.
A controlled Regional exercise should:
- freeze unrelated changes and record baseline;
- verify alternate readiness and capacity;
- block/withdraw primary traffic through a reversible fault injection;
- observe health detection and traffic steering;
- prove reads, writes, identity, asynchronous processing, and security logging;
- measure RTO, data lag/loss against RPO, error rate, and client recovery;
- operate in recovery long enough to expose hidden dependencies;
- restore primary capability without immediately shifting users;
- reconcile and resynchronize data;
- canary failback, then gradually restore traffic.
Failback is a separate migration, not an automatic undo. The former primary may contain stale data, old messages, certificates, or configuration. Define source of truth and conflict reconciliation before restoring bidirectional writes.
Cost and sustainability
Model:
- standby compute, databases, caches, storage, load balancers, NAT, firewalls, endpoints, TGW/Cloud WAN;
- inter-Region application and replication transfer in each charged direction;
- TGW peering/Cloud WAN data processing and attachments;
- Global Accelerator fixed and data-transfer premium, Route 53 queries/health checks, CloudFront;
- DX/VPN paths, data transfer out, cross-AZ paths, PrivateLink processing;
- logs, metrics, traces, replicated archives, security analysis, backups, KMS/secrets; and
- game days, operational staffing, licenses, support, and compliance.
Estimate normal, replication, failover, and failback traffic separately. During recovery, inter-Region reads or bulk resynchronization may dominate. Active/active doubles more than compute: it multiplies operational and testing scope. Choose the least complex pattern that meets measured business requirements.
Three-Region architecture workbook
Design Region A active, Region B warm standby, and Region C backup/pilot-light for a public API plus hybrid administration. Then compare an active/active alternative.
For at least 20 flows record:
| Field | Required evidence |
|---|---|
| Business | operation, RTO/RPO, capacity, residency |
| Entry | Route 53/GA/CloudFront, health, TTL/dial, TLS |
| Addressing | IPv4/IPv6 CIDRs, IPAM pool, ASN, overlap proof |
| Regional path | ingress, DNS, inspection, app, data, return |
| Inter-Region | TGW/Cloud WAN/service path and both routes |
| Data | authority, replication, lag, conflict and promotion |
| Dependency | Regional/global/external and isolation behavior |
| Failure | injected event, expected steering/reset/degradation |
| Evidence | metrics, logs, probe, transaction and audit |
| Cost | normal/failover transfer and standing owner |
Submit dependency inventory, CIDR/ASN plan, topology alternatives, 20-flow matrix, DNS/accelerator decision, Regional Resolver design, data-readiness contract, security/control-plane inventory, quota/capacity model, cost forecast, 12 game-day tests, failover/failback runbook, RACI, and proof that no resource changed.
Knowledge check
- Does a healthy alternate load balancer prove Regional readiness?
No. Data, identity, dependencies, capacity, security, and transactions must work.
- Do TGW peering routes propagate dynamically?
No. Peering uses explicit static routes on both Regional TGWs.
- Does DNS failover move existing TCP sessions?
No. Clients must reconnect, and cached DNS may delay new selection.
- Why can active/active be less safe?
Shared faults and write conflicts can affect every Region simultaneously.
- Why deploy local egress and DNS?
Routing those through another Region creates a hidden dependency during isolation.
- Is failback simply reversing the traffic control?
No. Data/configuration must be reconciled and the former primary requalified first.
Lesson acceptance
The lesson is complete only when the learner can:
- derive network design from RTO, RPO, capacity, residency, and failover authority;
- choose a recovery pattern and reject unjustified active/active complexity;
- create a unique IPv4/IPv6/IPAM/ASN plan for three Regions;
- choose TGW peering, Cloud WAN, or service-level connectivity by requirement;
- trace client, Regional, inter-Region, hybrid, DNS, security, data, and return paths;
- compare Route 53, Global Accelerator, and CloudFront behavior accurately;
- prove data and dependency readiness before traffic shift;
- separate pre-provisioned data-plane continuity from control-plane changes;
- diagnose partial and Regional failures using evidence;
- execute safe game-day, failover, reconciliation, and failback plans;
- calculate standing, replication, incident, and resynchronization cost; and
- attest that no staging or production AWS resource changed.