Lesson 275 · AWS Learning Path

AWS 275: Multi-Region network architecture

· Published · 15 min read

Labelled process diagram for AWS 275: Global client or hybrid source to Regional entry and route domain to Regional application dependencies to Health-directed response and recovery evidence, with decision, proof and...

Why this lesson matters

Deploying a second copy of a VPC does not create a recoverable multi-Region application. Clients need a healthy entry point, networks need unique routes, services need Regional dependencies, data needs a consistency and conflict model, identities and secrets must be available, and operators need a recovery process that still works when a control plane is impaired.

Multi-Region architecture buys a larger fault-isolation boundary only when the workload can run independently in another Region. It also introduces inter-Region transfer, replication lag, static routes or global policy, certificate/DNS coordination, security duplication, operational drift, and difficult failure modes. This lesson designs the network as one part of an application recovery system rather than treating global routing as disaster recovery.

Outcomes

By the end, you can:

  • derive network requirements from business impact, recovery time objective (RTO), and recovery point objective (RPO);
  • distinguish backup/restore, pilot light, warm standby, active/passive, and active/active;
  • create a non-overlapping Regional IPv4/IPv6 address and ASN plan with IPAM;
  • compare TGW inter-Region peering, AWS Cloud WAN, private endpoints, and application-layer patterns;
  • trace Regional, inter-Region, hybrid, DNS, ingress, egress, inspection, and return paths;
  • choose Route 53, Global Accelerator, CloudFront, or application discovery by protocol and failure behavior;
  • design Regional DNS, Resolver endpoints/rules, private hosted zones, and service discovery;
  • separate data-plane continuity from control-plane automation;
  • model data residency, encryption, identity, secrets, logging, and security policy;
  • test Region isolation, partial failure, stale health, replication lag, and failback; and
  • calculate inter-Region network, replication, endpoint, inspection, and standby cost.

Begin with the business outcome

Ask what business operation must continue, for whom, at what reduced capacity, after which failures, and with how much data loss. Then define:

TermQuestion
RTOHow long may the business capability remain unavailable?
RPOHow much committed data may be lost or require reconciliation?
Recovery capacityWhat percentage of normal traffic must the alternate Region serve?
Isolation scopeAZ, Region, dependency, operator error, identity, or data corruption?
ConsistencyCan clients read stale data or create conflicting writes?
ResidencyIn which jurisdictions may data, logs, backups, keys, and support access exist?
Failover authorityWho or what declares disaster and shifts traffic?
FailbackHow are data and traffic safely returned after recovery?

A 60-second DNS or accelerator failover cannot meet a 60-second RTO if the database needs two hours to restore. A zero-RPO requirement cannot be claimed with asynchronous replication that has measurable lag. The network design must use the application's realistic readiness.

Recovery patterns

PatternAlternate Region stateTypical tradeoff
Backup and restoreinfrastructure/data restored after eventlowest standing cost, longest RTO/RPO
Pilot lightcritical data/core services runningfaster recovery, substantial scale/configuration work remains
Warm standbycomplete reduced-capacity stackshorter RTO, ongoing capacity and drift management
Active/passivesecondary ready but not serving normal writessimpler data ownership, idle/standby cost
Active/active readsreads served in many Regions, writes controlledlatency gains with clearer write authority
Active/active writesseveral Regions accept writesfastest local writes, hardest conflict/consistency model

Active/active is not automatically more resilient. A bad deployment, corrupted replicated data, global credential, DNS policy, or shared dependency can fail every Region at once. Sometimes isolated active/passive plus tested promotion provides better safety.

Classify every dependency as Regional, global data plane, global control plane, external, or organizational. Record whether it is replicated, cached, reconstructed, replaced, or intentionally unavailable during recovery.

Addressing and IPAM

Allocate unique summarizable CIDRs per Region, environment, and trust zone. For example:

RegionAggregateProductionNonproductionShared
Region A10.64.0.0/1210.64.0.0/1410.68.0.0/1410.72.0.0/16
Region B10.80.0.0/1210.80.0.0/1410.84.0.0/1410.88.0.0/16
Region C10.96.0.0/1210.96.0.0/1410.100.0.0/1410.104.0.0/16

These are planning examples, not allocations to copy blindly. Check corporate, partner, acquisition, VPN, Direct Connect, container/pod, service, link-local, multicast, and IPv6 ranges. Summaries must not include prefixes routed somewhere else.

VPC IP Address Manager (IPAM) uses scopes, pools, and allocations. Build top-level private pools, Regional child pools with the correct locale, then environment/application pools. An IPAM pool allocates only in its locale. Select operating Regions deliberately, share pools with AWS RAM, enforce allocation tags/netmask ranges, and monitor overlap/compliance. Separate scopes can reuse addresses only when networks are guaranteed never to connect; that assumption often expires after mergers or analytics integration.

Plan IPv6 globally as well as IPv4. Separate IPv6 routing, egress, firewall, load balancer, DNS AAAA, client preference, and failure behavior. An alternate Region that works only over IPv4 is not dual-stack recovery.

Use unique TGW autonomous system numbers where possible. Although TGW peering uses static routes today, unique ASNs avoid ambiguity and support future/hybrid evolution.

Inter-Region connectivity choices

Transit Gateway peering

Peer Regional TGWs over the AWS global network. Peering supports static routes only: routes do not propagate across the peering attachment, and peering does not provide ECMP. Each relevant source-associated TGW route table needs a static route to the remote Regional summary, and the peer needs return routes. VPC subnet route tables still need their TGW routes.

Advantages include explicit segmentation and straightforward Regional hubs. Risks include route growth, manual/static lifecycle, hidden one-sided routes, summary mistakes, and operational scaling. TGW peering also does not make Route 53 Resolver in one Region transparently resolve names across the peer; design DNS endpoints/rules per Region.

AWS Cloud WAN

Cloud WAN provides a policy-managed global core network with segments, attachment policies, and Regional core network edges. It is useful when many Regions, sites, and SD-WAN attachments need consistent global segmentation and centralized policy. It does not eliminate VPC route tables, application recovery, DNS, data replication, or policy testing. A core network policy error can have broad blast radius, so version, review, simulate, and roll out policy changes.

Application and service-level connectivity

Not every workload needs routed inter-Region VPC connectivity. Public/edge APIs with TLS and identity, PrivateLink cross-Region endpoint services, asynchronous queues/events, database replication, object replication, or VPC Lattice patterns can expose only required services. Narrow service connectivity reduces route blast radius and overlap concerns.

RequirementStarting choice
Few Regional hubs and explicit routesTGW peering
Many Regions/sites with global segmentation policyCloud WAN
One private provider servicePrivateLink/cross-Region service pattern
Loose application integrationqueue/event/API/object replication
No inter-Region runtime dependencyindependent Regional stacks

Avoid circular runtime dependencies where Region A's application requires Region B's identity, DNS, firewall, logging, or database and Region B requires A. During partition, both can fail.

Trace three independent network planes

Client ingress

Internet or corporate clients must reach the selected healthy Region. Trace public DNS/accelerator, edge network, WAF/firewall, Regional load balancer/API, application, and return. Preserve client IP requirements, TLS name/certificate, IPv4/IPv6, sticky sessions, and long-lived connection behavior.

Regional east-west and egress

Each Region should have local DNS, inspection, NAT/egress, endpoints, and dependencies needed during isolation. Sending Region B's internet egress or logs through Region A creates a hidden Regional dependency. If policy deliberately centralizes something, state the failure mode and reduced service.

Inter-Region/hybrid

Trace every TGW/Cloud WAN route, attachment, inspection point, DX/VPN path, customer route, and reverse lookup. On-premises can advertise/receive multiple Regional prefixes, but BGP preference and asymmetric return paths need proof. A Region failover may overload VPN backup or customer firewalls sized only for normal traffic.

For every critical flow write both directions, including DNS and control calls. Use the AWS271 route-proof method and AWS274 inspection method.

Route 53 versus Global Accelerator

Route 53 returns DNS answers based on routing policy and health evaluation. Policies include failover, latency, geolocation, geoproximity, weighted, and multivalue behavior. Clients and intermediate resolvers cache answers for TTL; existing connections do not move merely because DNS changes. Health checks must measure the business-critical path, not only a load balancer listener.

Global Accelerator provides static anycast IP addresses and sends TCP/UDP traffic over the AWS global network to Regional endpoints. A standard accelerator chooses based on client location, endpoint health, traffic dials, and endpoint weights. New connections can move quickly when health changes; existing TCP sessions can still fail and reconnect. If no endpoints are healthy, Global Accelerator can route to all endpoints, so health and application failure behavior must be understood.

NeedRoute 53Global Accelerator
DNS-based web or service selectionstrong fitcan complement
static global IP allowlistingnot provided by DNS alonestrong fit
TCP/UDP non-HTTP accelerationDNS only selects addressstrong fit
fine geographic DNS policymany policy typesproximity/health endpoint groups
shift new traffic without waiting full DNS cacheconstrained by resolver/client TTLtraffic dial/health acts at accelerator
existing connection migrationnono transparent session migration

CloudFront is often the global entry for cacheable/content and HTTP workloads, with origin failover and WAF. API Gateway, ALB, and application-layer routers may fit other cases. Select based on protocol, static-IP need, caching, WAF, origin behavior, source preservation, TLS, and cost.

Global Accelerator traffic dial applies to traffic already directed to an endpoint group, not a percentage of all global traffic. Endpoint weights divide traffic within a group. Source-IP affinity can improve stickiness but can create uneven distribution and does not solve cross-Region session/data state.

Health, traffic shifting, and data readiness

A health signal must represent whether the Region can complete a safe transaction. Combine infrastructure and synthetic application checks carefully. A health check that writes production data may cause side effects; a shallow /health may stay green while database, identity, DNS, or queue dependency is broken.

Use readiness gates before traffic shift:

  • infrastructure and capacity deployed;
  • latest approved application/configuration version;
  • secrets, certificates, keys, and feature flags available;
  • data replication lag within RPO and promotion status known;
  • local DNS, endpoints, egress, inspection, logging, and identity working;
  • quotas and scaling tested for recovery load;
  • synthetic read and controlled write/rollback pass; and
  • incident commander has explicit go/no-go evidence.

Use weighted DNS or accelerator dials for a canary, observe errors/latency/replication, then increase gradually. Automatic failover is appropriate only when signals are reliable and failover cannot cause split-brain writes. Otherwise automate evidence collection and require controlled approval.

Data, state, and consistency

Network reachability cannot make a secondary database writable safely. For each store record replication type, direction, lag, consistency, conflict resolution, failover/promotion, endpoint discovery, backup independence, encryption key, and failback/resynchronization.

Read-local/write-primary patterns reduce write conflicts but make the primary Region a dependency. Global databases can support multi-Region reads or writes with service-specific semantics. S3 cross-Region replication is asynchronous and existing-object replication/versioning/deletion behavior must be designed. Queues, caches, search indexes, file systems, and object stores all have different replication guarantees.

Sessions should be portable, reconstructible, or deliberately reauthenticated. Do not rely on load-balancer stickiness to preserve a session after Regional failover. Idempotency keys, globally unique identifiers, clocks, duplicate event handling, and conflict rules belong in the recovery design.

Data residency includes replicas, backups, snapshots, logs, traces, DNS/query records, security findings, and support access. Encryption in transit and at rest does not by itself satisfy location restrictions.

DNS and service discovery

Create Regional Route 53 Resolver inbound/outbound endpoints and rules when each Region must operate independently. Share/associate private hosted zones and Profiles according to Region/account behavior. Test split-view DNS, negative caching, stale records, and private-zone conflicts.

Use Regional service names such as api.region-a.internal.example plus a controlled global name when helpful. A global private name should not direct clients to a Region whose data is not ready. Service discovery records need health and lifecycle semantics; DNS registration alone is not application readiness.

TGW peering does not provide cross-Region Amazon-provided DNS resolution. Forwarding DNS across a peering link to one central Region recreates a dependency. Prefer local Resolver paths with deliberate cross-Region rules only where required.

Inspection and security independence

Deploy Regional inspection, NAT, DNS filtering, VPC endpoints, WAF, logging buffers, and security controls needed for standalone operation. Replicate policies through versioned automation, but permit controlled Regional rollout to limit global blast radius.

Central policy services such as Firewall Manager or Cloud WAN simplify governance but their data-plane and control-plane failure behavior must be documented. Existing Regional resources may continue processing while configuration APIs are unavailable. Do not make failover require creating a new firewall, endpoint, certificate, or IAM role during an incident if the RTO cannot tolerate that dependency.

Use separate break-glass roles, tested access paths, hardware MFA/process controls, and audit destinations. Replicate secrets only to approved Regions using service-supported mechanisms. KMS keys are Regional resources even when related multi-Region key features are used; policies, grants, aliases, and service integration still require proof.

Control plane versus data plane

The data plane processes established service traffic. The control plane creates or changes DNS records, routes, accelerators, scaling, certificates, and policies. A resilient design pre-provisions its recovery data plane so routine failover mainly changes a small, tested steering control.

Avoid requiring a long chain of control-plane calls during failure. Examples include creating a VPC, requesting quota, restoring every secret, provisioning endpoints, issuing certificates, and then changing DNS. Pre-create and continuously validate what the RTO needs.

Route 53 is a global service whose control plane has Regional hosting details; its globally distributed authoritative data plane is separate. Do not equate one console/API impairment with DNS data-plane failure. Cache emergency runbooks and approved infrastructure definitions outside the affected operational path.

Read-only evidence collection

Inventory every enabled Region explicitly:

aws sts get-caller-identity --query Arn --output text
aws ec2 describe-regions +  --query 'Regions[].{Region:RegionName,OptIn:OptInStatus}' --output table

for region in ap-south-1 ap-southeast-1 eu-west-1; do
  aws ec2 describe-transit-gateways --region "$region" --output json
  aws ec2 describe-transit-gateway-peering-attachments +    --region "$region" --output json
  aws ec2 describe-vpcs --region "$region" +    --query 'Vpcs[].{Id:VpcId,Cidrs:CidrBlockAssociationSet[].CidrBlock}' +    --output json
done

Global and multi-Region services:

aws route53 list-health-checks --output json
aws globalaccelerator list-accelerators --region us-west-2 --output json
aws networkmanager list-core-networks --region us-west-2 --output json
aws ec2 describe-ipams --region ap-south-1 --output json
aws ec2 describe-ipam-pools --region ap-south-1 --output json

Global Accelerator API calls use us-west-2. Network Manager/Cloud WAN command Regions and resource scope must be confirmed for the deployed design. Also inspect TGW routes/associations/propagations, VPC route tables, Cloud WAN policy versions/routes, Route 53 records/health status, Resolver, WAF/firewalls, load balancers, DX/VPN/BGP, quotas, replication metrics, CloudTrail, Config, CUR, and application SLOs. Redact identities, IP plans, domains, routes, and recovery topology.

Failure diagnosis

SymptomFirst proofCommon cause
DNS selects unhealthy Regionhealth-check chain and TTLshallow/stale health or cached answer
Accelerator stays in Regionendpoint/endpoint-group healthhealthy LB but broken dependency
New Region reachable, writes faildatabase promotion/lagnetwork shifted before data readiness
One-way inter-Region flowstatic TGW routes both sidesmissing return route
Remote hostname failsRegional Resolver designassumed DNS across TGW peering
On-premises uses old RegionBGP advertisements/preferencesroute not withdrawn or customer preference
Recovery overloadsquotas/capacity metricsstandby never load-tested
Only IPv6 clients failAAAA/IPv6 route/securitypartial dual-stack deployment
Security logs disappearRegional destination/dependencylogs depended on failed Region
Failback duplicates datareplication/conflict evidenceno reconciliation/idempotency process

Start at client resolution/accelerator selection, then Regional ingress, application dependencies, data authority, inter-Region and hybrid routes, security path, and return. Correlate UTC evidence from both Regions. Do not change DNS and routes simultaneously without knowing which change corrected the fault.

Game days, failover, and failback

Test isolated dependency failure before whole-Region simulation. Include load balancer health, DNS Resolver, firewall/NAT, TGW/Cloud WAN route, DX, identity, secrets, database lag, logging destination, control-plane API, and operator access.

A controlled Regional exercise should:

  1. freeze unrelated changes and record baseline;
  2. verify alternate readiness and capacity;
  3. block/withdraw primary traffic through a reversible fault injection;
  4. observe health detection and traffic steering;
  5. prove reads, writes, identity, asynchronous processing, and security logging;
  6. measure RTO, data lag/loss against RPO, error rate, and client recovery;
  7. operate in recovery long enough to expose hidden dependencies;
  8. restore primary capability without immediately shifting users;
  9. reconcile and resynchronize data;
  10. canary failback, then gradually restore traffic.

Failback is a separate migration, not an automatic undo. The former primary may contain stale data, old messages, certificates, or configuration. Define source of truth and conflict reconciliation before restoring bidirectional writes.

Cost and sustainability

Model:

  • standby compute, databases, caches, storage, load balancers, NAT, firewalls, endpoints, TGW/Cloud WAN;
  • inter-Region application and replication transfer in each charged direction;
  • TGW peering/Cloud WAN data processing and attachments;
  • Global Accelerator fixed and data-transfer premium, Route 53 queries/health checks, CloudFront;
  • DX/VPN paths, data transfer out, cross-AZ paths, PrivateLink processing;
  • logs, metrics, traces, replicated archives, security analysis, backups, KMS/secrets; and
  • game days, operational staffing, licenses, support, and compliance.

Estimate normal, replication, failover, and failback traffic separately. During recovery, inter-Region reads or bulk resynchronization may dominate. Active/active doubles more than compute: it multiplies operational and testing scope. Choose the least complex pattern that meets measured business requirements.

Three-Region architecture workbook

Design Region A active, Region B warm standby, and Region C backup/pilot-light for a public API plus hybrid administration. Then compare an active/active alternative.

For at least 20 flows record:

FieldRequired evidence
Businessoperation, RTO/RPO, capacity, residency
EntryRoute 53/GA/CloudFront, health, TTL/dial, TLS
AddressingIPv4/IPv6 CIDRs, IPAM pool, ASN, overlap proof
Regional pathingress, DNS, inspection, app, data, return
Inter-RegionTGW/Cloud WAN/service path and both routes
Dataauthority, replication, lag, conflict and promotion
DependencyRegional/global/external and isolation behavior
Failureinjected event, expected steering/reset/degradation
Evidencemetrics, logs, probe, transaction and audit
Costnormal/failover transfer and standing owner

Submit dependency inventory, CIDR/ASN plan, topology alternatives, 20-flow matrix, DNS/accelerator decision, Regional Resolver design, data-readiness contract, security/control-plane inventory, quota/capacity model, cost forecast, 12 game-day tests, failover/failback runbook, RACI, and proof that no resource changed.

Knowledge check

  1. Does a healthy alternate load balancer prove Regional readiness?

No. Data, identity, dependencies, capacity, security, and transactions must work.

  1. Do TGW peering routes propagate dynamically?

No. Peering uses explicit static routes on both Regional TGWs.

  1. Does DNS failover move existing TCP sessions?

No. Clients must reconnect, and cached DNS may delay new selection.

  1. Why can active/active be less safe?

Shared faults and write conflicts can affect every Region simultaneously.

  1. Why deploy local egress and DNS?

Routing those through another Region creates a hidden dependency during isolation.

  1. Is failback simply reversing the traffic control?

No. Data/configuration must be reconciled and the former primary requalified first.

Lesson acceptance

The lesson is complete only when the learner can:

  • derive network design from RTO, RPO, capacity, residency, and failover authority;
  • choose a recovery pattern and reject unjustified active/active complexity;
  • create a unique IPv4/IPv6/IPAM/ASN plan for three Regions;
  • choose TGW peering, Cloud WAN, or service-level connectivity by requirement;
  • trace client, Regional, inter-Region, hybrid, DNS, security, data, and return paths;
  • compare Route 53, Global Accelerator, and CloudFront behavior accurately;
  • prove data and dependency readiness before traffic shift;
  • separate pre-provisioned data-plane continuity from control-plane changes;
  • diagnose partial and Regional failures using evidence;
  • execute safe game-day, failover, reconciliation, and failback plans;
  • calculate standing, replication, incident, and resynchronization cost; and
  • attest that no staging or production AWS resource changed.

Official sources

Advertisement