Lesson 330 · AWS Learning Path

AWS 330: Multi-account and multi-Region design

· Published · 8 min read

Labelled process diagram for AWS 330: Organization and regional requirements to Account and Region architecture to Workload deployment and data paths to Isolation, failover, operations, and cost evidence, with...

Why this lesson matters

Accounts isolate ownership, policy, quota, billing, and blast radius. Regions isolate infrastructure and support residency, proximity, sovereignty, and regional-disaster requirements. Combining both can improve enterprise resilience, but also multiplies identity, network, data, deployment, observability, security, recovery, and cost complexity.

Most workloads can meet availability needs with multiple Availability Zones in one Region. Use multiple Regions only for an explicit business, regulatory, proximity, or recovery requirement. A copied stack is not a multi-Region system unless data, traffic, dependencies, authority, tests, and failback work.

Learning outcomes

By the end, you can:

  • derive account and Region placement from requirements;
  • distinguish Multi-AZ high availability from multi-Region recovery/availability;
  • choose backup/restore, pilot light, warm standby, or active-active;
  • design identity, network, DNS/traffic, data, keys, deployment, and telemetry;
  • identify global and regional control/data-plane dependencies;
  • design data sovereignty and cross-border evidence;
  • plan failover, fencing, reconciliation, and failback;
  • model the cost and operations of a 40-account, three-Region platform.

1. Start with separate dimensions

Do not use accounts and Regions as interchangeable isolation tools.

DimensionAccount boundaryRegion boundary
Primary purposeOwnership, policy, billing, quota, blast radiusLocation, latency, residency, service/failure isolation
GovernanceOrganizations, OUs, SCP/RCP, identity assignmentsRegion enablement, service availability, residency policy
NetworkVPC sharing/peering/TGW/Cloud WAN and policiesInter-Region routing, encryption, bandwidth, transfer
DataCross-account policy and key ownershipReplication, consistency, conflict, RPO, sovereignty
RecoverySeparate credentials/resources/evidenceRegional disaster strategy and traffic failover

Map each workload by business owner, environment, risk class, data jurisdiction, user location, RTO/RPO, availability, cost center, and support team. Then derive account and Region placement.

2. Account architecture

Use management, log archive, audit/security, network, shared services, and workload accounts as justified in AWS324. Separate production and nonproduction. Avoid workloads in the management account. Apply controls to OUs based on stable policy needs and use delegated administrators for supported services.

For each account record owner, purpose, OU, data class, Regions, identity groups, network pattern, cost center, security/backup integrations, support, and lifecycle state. Account vending and closure must work in every required Region.

Account separation does not automatically isolate shared networks, pipelines, identity, DNS, keys, or security tooling. Map cross-account trust and shared-service blast radius explicitly.

3. Region selection

Filter candidates by legal/residency, required service/feature, connectivity, customer latency, partner dependencies, support, quota, and recovery separation. Environmental and cost factors compare remaining candidates.

Region codes in diagrams are not proof. Verify service and feature availability, opt-in state, quotas, instance types, certificate/key behavior, partner connectivity, and organizational controls. Data sovereignty includes logs, backups, support access, analytics, telemetry, and security findings, not just the primary database.

4. Multi-AZ before multi-Region

Use independent Availability Zones to protect against common infrastructure failures. Ensure every tier distributes capacity and does not depend on one NAT gateway, mount target, subnet, database instance, endpoint, or self-managed appliance.

Multi-Region adds asynchronous replication, routing, duplicate infrastructure, control-plane operations, consistency decisions, and failover/failback. AWS guidance notes it is usually unnecessary for most workloads. Justify it with a requirement that Multi-AZ cannot meet.

5. Recovery strategy

StrategyRecovery Region before eventTypical trade-off
Backup and restoreBackups and deployable artifactsLowest steady cost; longest RTO and control-plane dependence
Pilot lightCore data/services active; compute mostly absentLower cost; scale/deploy steps during recovery
Warm standbyScaled-down complete workloadFaster recovery; continuous duplicate operation
Active-activeBoth Regions serve productionLowest potential RTO; highest data/conflict/operation complexity

Do not copy generic RTO/RPO values. Measure deployment, restore, scale, DNS/traffic, cache/session, dependency, validation, and decision time. Recovery objectives must be less than maximum tolerable disruption and data loss.

Active-active requires write ownership, conflict handling, global uniqueness, idempotency, session behavior, partition strategy, and reconciliation. Replication does not protect against corruption or malicious deletion without point-in-time or isolated backup.

6. Global and regional dependency map

Classify every dependency:

  • globally controlled/global data plane;
  • globally controlled/regional data plane;
  • regional control and data plane;
  • external enterprise or third party.

Ask whether existing resources continue to operate if a control plane is impaired. Avoid requiring resource creation, policy changes, secret retrieval, or DNS edits that were never tested during recovery. Pre-provision and validate what the RTO demands.

Identity Center, Organizations, IAM, Route 53, CloudFront, Global Accelerator, artifact repositories, CI/CD, secrets, KMS, DNS, monitoring, and ticket/communication systems each have distinct scope. Do not call them all simply “global.” Verify documented behavior and design fallback.

7. Network and DNS/traffic design

Plan nonoverlapping address space with IPAM, regional VPCs, multiple AZs, ingress, egress, endpoints, DNS, hybrid connectivity, inspection, and routing. Decide TGW, Cloud WAN, peering, or service exposure by scale and route/security requirements. Inter-Region network paths need bandwidth, encryption, MTU, route symmetry, failure, and transfer-cost analysis.

Route 53 health-check/failover, latency, geolocation, or weighted policies and Global Accelerator solve different traffic needs. Health must represent the user journey and regional readiness, not only one healthy endpoint. DNS TTL and client caching affect recovery. Global Accelerator uses anycast static IPs and health-based endpoint routing, but application/data readiness remains the customer's job.

Prevent split brain. Define who can declare a Region unhealthy, stop writes, fence the old writer, change traffic, and reverse the decision.

8. Data architecture

For every store define system of record, replication direction, synchronous/asynchronous behavior, measured lag, consistency, write topology, keys, schema, backup, PITR, conflict handling, failover, failback, and residency.

Examples include S3 replication, DynamoDB global tables, Aurora Global Database, RDS cross-Region replicas, OpenSearch cross-cluster replication, streaming replication, or application-level transfer. Features and semantics differ; do not infer one from another.

Multi-Region KMS keys provide related key material across Regions for supported designs, but each regional key has policy, grants, state, and operation. They do not replicate application data or authorize identities automatically. Ensure recovery operators can use keys without weakening separation of duties.

9. Deployment and configuration

Use infrastructure as code and a controlled promotion process across accounts and Regions. Keep templates consistent while allowing explicit regional parameters. Store artifacts where recovery can access them. Validate AMIs/images, packages, certificates, secrets, feature flags, quotas, and dependencies in each Region.

Avoid deploying all Regions simultaneously. Use waves/cells, automated tests, canaries, and stop conditions. Configuration drift can make standby recovery fail even when the main stack is healthy. Continuously compare required regional capability, not only stack status.

10. Security and observability

Aggregate security and operational evidence while retaining regional survivability and legal boundaries. Ensure CloudTrail, Config, GuardDuty, Security Hub, logs, metrics, traces, flow logs, backup status, and cost are enabled in intended Regions/accounts and delegated correctly.

Centralization can create a blind spot during network or destination failure. Define local buffering/retention, cross-account write permissions, immutable evidence, and alternate access. Test security response and break-glass separately from workload failover.

Dashboards need per-Region and global user outcomes, replication lag, queue/backlog, data conflict, traffic distribution, regional saturation, quota, synthetic journey, and recovery readiness. A green secondary with no representative test traffic may be silently broken.

11. Failover and failback

A runbook includes incident criteria, authority, communication, evidence freeze, dependency checks, data/RPO decision, write fencing, infrastructure activation, capacity scale, secrets/keys, traffic shift, validation, observation, and rollback/forward recovery.

Failback is a separate migration. Determine new source of truth, replication direction, data divergence, reconciliation, rehydration, traffic stages, and acceptance. Do not automatically fail back simply because the primary Region returns.

Run game days for AZ loss, Region isolation, bad deployment, data corruption, identity failure, DNS issue, dependency loss, and unavailable operator. Record actual RTO/RPO and repair gaps.

12. Read-only discovery

aws sts get-caller-identity --query Arn --output text
aws organizations list-accounts --output table
aws ec2 describe-regions --all-regions --output table
aws route53 list-health-checks --output table
aws globalaccelerator list-accelerators --region us-west-2 --output table
aws backup list-backup-vaults --region ap-south-1 --output table
aws backup list-backup-vaults --region ap-southeast-1 --output table

Permissions and service endpoints vary. Inventory is not proof that failover works.

13. Guided workshop: 40 accounts, three Regions

Design for regulated payments in India, customer services in Singapore, and DR/reporting in another approved Region. Produce:

  1. business, legal, latency, and recovery requirements;
  2. account/OU/owner map for 40 accounts;
  3. workload-to-account-to-Region placement matrix;
  4. service/feature/quota Region evidence;
  5. Multi-AZ versus multi-Region justification per workload;
  6. recovery strategy and measured RTO/RPO budget;
  7. global/regional/external dependency register;
  8. IP, routing, DNS, and traffic architecture;
  9. identity, delegated administration, and break-glass design;
  10. data system-of-record/replication/conflict map;
  11. key, secret, certificate, and artifact availability plan;
  12. infrastructure/configuration deployment waves;
  13. security and observability aggregation with local survival;
  14. sovereignty and cross-border data-flow register;
  15. failover runbook with fencing and authority;
  16. failback and reconciliation runbook;
  17. ten game-day cases and acceptance evidence;
  18. normal, standby, recovery, transfer, and operations cost model.

14. Troubleshooting

SymptomInvestigate
Healthy endpoint receives no trafficHealth check, routing policy, TTL/cache, GA endpoint weight
Failover app starts but login failsIdP, redirect URI, certificate, secret, key, network dependency
Secondary cannot accept loadQuota, scaling, warm capacity, downstream connection limit
Data is present but staleReplication lag/error, writer topology, checkpoint, conflict
Security has regional blind spotDelegated admin, Region enrollment, log destination, policy
Failback risks lossNew writes, replication direction, divergence, reconciliation

Cost and cleanup

This lesson creates no resources. Include duplicate capacity, replication requests/storage, inter-Region and internet transfer, TGW/Cloud WAN, NAT/inspection, DNS/Global Accelerator, logs/security, backup, support, game days, and staffing. Active-active is not justified merely because it is technically possible.

Knowledge check

  1. Does multi-account imply multi-Region? No; they solve different isolation and location requirements.
  2. Is multi-Region necessary for most workloads? No; Multi-AZ commonly meets availability requirements.
  3. Why is replication not backup? Corruption or deletion can replicate without PITR/isolated recovery.
  4. What prevents split brain? Explicit write ownership, fencing, authority, and tested traffic/data procedures.
  5. Why test failback? It requires data reconciliation and traffic migration in the opposite direction.

Lesson acceptance

Submit all 18 artifacts. Every account/Region placement must trace to a requirement; regional service, data, key, network, dependency, quota, and operator readiness must be evidenced; and failover/failback must prove fencing, reconciliation, actual RTO/RPO, user acceptance, and cost.

Official sources

Advertisement