AWS 262: Multi-account network architecture
Why this matters to an architect
An AWS account boundary does not create, prevent, or inspect network traffic. A secure multi-account design must deliberately join VPC ownership, addresses, routing domains, DNS, hybrid links, ingress/egress, private service access, inspection, identity, telemetry, capacity, cost, and recovery.
Centralizing everything can reduce duplicated appliances but creates shared blast radius, policy scale, inter-AZ/Region paths, and an operations bottleneck. Distributing everything can improve isolation and autonomy but multiplies endpoints, routes, cost, and configuration drift. The architect chooses per capability from measurable requirements - not from a fashionable hub diagram.
Outcomes
By the end, you can:
- derive network requirements and trust zones from workload flows;
- choose per-account VPCs, shared VPCs, or a portfolio hybrid;
- compare peering, Transit Gateway (TGW), and Cloud WAN;
- design route tables/segments for isolation, shared services, inspection, and hybrid paths;
- plan IPv4/IPv6 with delegated VPC IPAM and RAM-shared pools;
- design centralized or distributed DNS, endpoints, ingress, egress, and firewalls;
- preserve stateful-routing symmetry and multi-AZ/Region failure domains;
- design redundant Direct Connect/VPN and BGP policy;
- assign owner/participant responsibilities and automate safe changes;
- prove reachability, denied paths, DNS, performance, telemetry, cost, failover, and rollback.
Begin with flows and failure requirements
Inventory every required and forbidden path:
| Dimension | Questions that change design |
|---|---|
| source/destination | account, VPC/subnet, on-premises/site, internet, partner, AWS service, SaaS |
| protocol | IPv4/IPv6, TCP/UDP/ICMP, ports, DNS, multicast, MTU/fragmentation |
| trust/data | environment, tenant, data class, regulatory/residency, source identity |
| performance | latency, jitter, throughput, packets/flows, connections, DNS QPS |
| resilience | AZ/Region/site/device failures, RTO/RPO, degraded operation |
| inspection | stateful/stateless, TLS, IDS/IPS, source/destination preservation, fail-open/closed |
| ownership | requester, approver, route/DNS/firewall owner, workload responder |
| evidence | flow/firewall/DNS logs, route/config history, active probes, packet evidence |
| cost | hourly attachments/endpoints/NAT/firewall, processing, inter-AZ/Region/egress |
Draw the forward and return path hop by hop. Include DNS before connection establishment and identity/service policy before data flow. “Connected to TGW” is not an application-flow design.
Account and VPC ownership models
VPC per workload account
The workload account owns VPC/subnets/routes/security groups and application resources, while central IPAM, TGW/Cloud WAN, Resolver rules, firewall policy, and guardrails supply the platform. This gives strong ownership/failure boundaries and independent endpoint/security policy, but increases VPCs, attachments, address use, endpoint/NAT cost, and lifecycle automation.
Use when workloads need independent routing, sensitive isolation, delegated change cadence, or separate lifecycle. Standardize with account vending/IaC; do not allow arbitrary CIDRs and unmanaged internet gateways.
Shared VPC
The network account owns a VPC and shares selected subnets through RAM with participant accounts. Participants create supported application resources in those subnets, own their resources/security groups, and cannot manage owner/other-participant resources. The owner controls VPC CIDRs, subnets, route tables, NACLs, gateways, DHCP/DNS attributes, and VPC-level telemetry.
Benefits include fewer VPCs/attachments, efficient IPv4 use, simple same-VPC connectivity, and no inter-account data-transfer charge for same-AZ instance traffic. Tradeoffs include shared IP/routing/NACL/DNS failure domains, owner bottleneck, quota contention, reduced participant control, and unsupported resource types. Sharing separate subnets per participant improves NACL/routing separation but is not equivalent to a VPC boundary.
Use for cooperative teams with similar trust/connectivity/lifecycle requirements. Define IP allocation, AZs, route/NACL/security-group responsibility, flow-log access, quota, incident, participant removal, cost, and resource deletion order.
Hybrid portfolio
Common enterprise answer: dedicated VPCs for regulated/production/unique routing, shared VPCs for aligned platform or development portfolios, and TGW/Cloud WAN/PrivateLink between domains. Document why each portfolio shares a failure domain.
Choose the connectivity mechanism
| Mechanism | Good fit | Reject or constrain when |
|---|---|---|
| VPC peering | small number of direct, non-transitive VPC relationships | mesh scale, overlapping CIDRs, transitive hub/inspection needed |
| Transit Gateway | Regional hub-and-spoke, multiple route domains, VPC/VPN/DX/Connect, centralized services/inspection | global policy consistency is dominant or per-flow narrow service exposure is enough |
| Cloud WAN | policy-driven multi-Region/site core, global segments, attachment policy, service insertion, centralized visibility | small single-Region estate where operational/cost complexity is unjustified |
| PrivateLink | expose one service privately with producer/consumer isolation and no broad routing | general bidirectional network connectivity is required |
| VPC Lattice | service-network connectivity and policy for supported application services | arbitrary network protocols/routing or unsupported resources |
| shared VPC | participant resources need one owner-controlled VPC | strong VPC-level isolation/autonomy needed |
Do not select solely by maximum scale. Compare operational model, segmentation, propagation control, route convergence, multi-Region behavior, appliance insertion, IPv6, quotas, observability, cost, migration, and rollback.
Transit Gateway route-domain design
A TGW is Regional. It routes IPv4/IPv6 between attachments using TGW route tables. Each attachment associates with one route table and can propagate routes to multiple route tables. VPC subnet route tables still require routes to TGW, and destination/return routes must exist.
Do not leave every attachment in default association/propagation. Design route tables such as:
Prod attachment -> Prod ingress table -> approved Prod/shared/inspection paths
NonProd attachment -> NonProd table -> isolated NonProd/shared paths
Inspection attachment -> Inspection table -> spoke/hybrid return routes
Hybrid attachment -> Hybrid table -> approved summarized AWS prefixes
Quarantine attachment -> no propagation; explicit remediation services only
Static routes override propagated routes for the same destination; blackhole routes intentionally drop matching traffic. TGW peering requires static routes. VPC propagation advertises VPC CIDRs; BGP attachments introduce learned routes. Record association, propagation, static/blackhole, prefix ownership, and expected route evaluation.
TGW sharing through RAM lets participant accounts create VPC attachments; either side can delete an attachment. Approval, tags, route association/propagation, change ownership, and deletion alarms must be automated. A participant-created attachment must not automatically join a privileged default table.
Cloud WAN design
Cloud WAN creates a global/core network with Regional core network edges defined by a versioned JSON policy. Segments are isolated routing domains by default. Attachment-policy rules map attachments using attachment tags/metadata, in rule-number order; first match wins, and a miss can leave the attachment unassociated. Acceptance can be required.
The policy defines Regions, segments, sharing, routing, attachment policies, network-function groups, and service insertion. Policy versions and change sets support review/rollback. Current policy version 2025.11 is required for newer routing policies/BGP-community capabilities; verify feature/Region/attachment support rather than copying syntax blindly.
Use RAM to share the core network. The core owner controls policy, edges, acceptance, and global routing; attachment owners manage their approved attachments/tags. Because tags can place an attachment in a segment, protect tag mutation and default-deny unmatched/malformed cases.
Cloud WAN does not remove the need for VPC routes, DNS, address planning, firewall behavior, endpoint policy, hybrid redundancy, or workload security controls.
IP addressing and VPC IPAM
Create a hierarchy from enterprise space to environment/Region/portfolio pools. Delegate an IPAM member account through the supported IPAM Organizations workflow so its service-linked role can monitor organization use; merely registering generic trusted access incorrectly can miss required setup. Share pools through RAM to accounts/OUs.
Define allocation locale, allowed netmask, auto-import/discovery, required tags, provisioned CIDRs, utilization thresholds, reservation/growth, BYOIP/public space, release quarantine, and M&A overlap. Account vending requests CIDRs idempotently and records allocation/resource/account/Region owner.
IPv4 plans must include endpoints, TGW/appliance/DNS subnets, scaling, blue-green, DR, acquisitions, and on-premises. Avoid giant VPCs “for future use” and tiny subnets that block scaling. NAT is not an address-management strategy.
Plan IPv6 intentionally: dual-stack versus IPv6-only subnets, egress-only internet gateway, DNS/AAAA, NAT64/DNS64 where needed, security controls, logs/tools, hybrid routing, load balancer/service support, and applications. IPv6 removes IPv4 exhaustion, not segmentation or egress governance.
DNS architecture
Every VPC normally uses Route 53 VPC Resolver at its VPC-provided address. Do not configure workloads to send ordinary DNS directly to outbound Resolver endpoint IPs.
Central hybrid pattern:
- multi-AZ inbound endpoints let on-premises resolvers query approved Route 53 private namespaces;
- multi-AZ outbound endpoints forward matching domains to on-premises/partner resolvers;
- forwarding/system rules are shared through RAM and associated with spoke VPCs;
- private hosted zones or Route 53 Profiles/Global Resolver capabilities are associated/shared according to current support;
- DNS Firewall rule groups, query logs, and domain ownership add control/evidence.
Forwarding a rule does not require spoke-to-endpoint VPC routing because the service handles it, but target DNS servers must be reachable from outbound endpoints. Define conditional-forwarding precedence, most-specific matches, split horizon, delegation, overlapping zones, loops, NXDOMAIN, TTL/cache, DNSSEC, endpoint IP/capacity, and failover.
Unsharing/deleting a Resolver rule changes associated VPC behavior to remaining rules. Treat as a production dependency. Test AWS-to-on-premises, on-premises-to-AWS, denied domains, target/endpoints failure, and alternate path.
Private access to AWS and application services
Gateway endpoints provide private S3/DynamoDB routing with no endpoint hourly fee; route and endpoint/resource policies remain critical. Interface endpoints use PrivateLink ENIs, security groups, endpoint policy, per-AZ hourly and data-processing charges.
Distributed interface endpoints give each VPC independent availability/policy/blast radius but multiply cost and operations. Centralized endpoints reduce copies but add TGW/routing/DNS/data processing, policy-size/least-privilege complexity, and shared blast radius. Managed private DNS for an interface endpoint applies to its endpoint VPC; central consumers may require carefully managed private hosted-zone/profile/Resolver design. Compare total path cost and failure, not endpoint hourly price alone.
PrivateLink exposes a service without transitive routing or overlapping-CIDR concern and preserves a narrow trust boundary. It is often safer than connecting entire partner/tenant VPCs. Define endpoint-service permissions, acceptance, NLB/GWLB health, AZs, DNS, consumer endpoint policy/security groups, logging, quotas, and producer/consumer billing.
Internet ingress and egress
Ingress choices include CloudFront/WAF/Shield to regional ALB/API Gateway, Global Accelerator, public NLB/ALB, or controlled appliances. Prefer service-edge protection and application authentication over routing public traffic through a universal network hub. Define certificates/DNS, origin restriction, client IP, health, DDoS, IPv6, logs, failover, and cross-zone/Region cost.
For egress, compare distributed NAT gateway/firewall per VPC with centralized egress VPC through TGW/Cloud WAN. Centralized egress improves policy and public-IP control but creates shared capacity/failure, longer paths, TGW/firewall/NAT/inter-AZ processing, and source attribution concerns. Deploy NAT/firewall endpoints per used AZ and keep AZ-local routing where required; cross-AZ “HA” can increase cost and hide an undersized design.
Use gateway/interface endpoints to avoid internet/NAT where suitable. Domain filtering, proxying, Network Firewall, or third-party controls must align with encrypted DNS/TLS and application requirements. Decide fail-open/closed explicitly.
Centralized inspection and routing symmetry
AWS Network Firewall and stateful appliances require both directions of a flow to traverse the same firewall endpoint/state context. In a TGW inspection VPC, use dedicated TGW attachment and firewall subnets/route tables, enable TGW appliance mode on the inspection attachment, and route forward/return paths through the firewall. Network Firewall does not support asymmetric routing.
Inspect selected trust-zone crossings rather than every packet by default. Validate source/destination preservation, SNAT order, fragment/MTU, east-west/on-premises/internet flows, endpoint scaling, rule capacity, Suricata/stateless-stateful ordering, TLS limitations, logging, and bypass routes. A green firewall endpoint does not prove traffic traverses it.
Cloud WAN service insertion uses network-function groups and policy actions to redirect selected same/cross-segment paths; tags and policy become high-risk routing controls. Test the deployed policy and actual flow path.
Hybrid and multi-Region connectivity
Direct Connect needs redundant connections, routers, devices, and preferably locations according to the required resiliency model. Maximum resiliency uses separate connections terminating on separate devices in more than one location. VPN can provide backup or primary paths; test health and failover rather than assuming BGP handles everything.
Document Direct Connect gateway/TGW/core associations, virtual interfaces, ASNs, BGP communities/local preference/AS path/MED, allowed prefixes, route limits, summarization, default-route policy, MTU, MACsec where applicable, encryption requirements, and maintenance ownership. Use the Direct Connect failover test and active network probes.
TGW peering, Cloud WAN, inter-Region VPC peering, and service-specific replication solve different needs. Build Regional failure domains: do not send all regions' egress, DNS, inspection, or endpoints through one Region unless the business accepts that dependency and latency/data-transfer cost. Define route convergence, stateful-session loss, DNS failover, control-plane access, and degraded local operation.
Security layers and data perimeter
Use routing/segments to limit reachability, NACLs for stateless subnet controls where justified, security groups for stateful workload ENI policy, endpoint/resource policies for service access, Network Firewall/GWLB for inspection, WAF for HTTP, and IAM/SCP/RCP for API/resource authorization. Network reachability is not authorization; IAM authorization is not network reachability.
Security-group referencing has topology/Region/service constraints; verify current TGW/Cloud WAN and cross-account support. Prefix lists simplify approved CIDRs but require version/change governance. Data perimeter controls may use VPC endpoints and aws:SourceVpce/aws:SourceVpc or organization context where supported; test service-to-service exceptions and break glass.
Observability and evidence
Collect VPC/TGW flow logs, Network Firewall flow/alert logs, DNS query logs, ELB/CloudFront/WAF logs, Direct Connect/VPN/TGW/Cloud WAN metrics/events, IPAM compliance/utilization, Reachability Analyzer/Network Access Analyzer findings, CloudTrail changes, AWS Config, and active synthetic probes.
Flow logs are sampled/aggregated records, not packet captures, and ACCEPT does not prove application success. Correlate DNS, routes, security groups/NACLs, firewall/NAT, load balancer, host, and application evidence with timestamps. Protect central logs as in AWS260.
Safe automation and change
Keep IPAM, VPC/subnet, RAM shares, attachments, route tables, Cloud WAN policies, DNS, endpoints, firewall, hybrid configuration, and monitoring in versioned pipelines with one authoritative controller per resource. Validate schemas and route impact, model forward/return paths, run policy/reachability tests, deploy to a network test account/segment, canary attachments/routes, monitor, then batch.
Prevent route leaks with approved-prefix registries, summarization, max-prefix/route quotas, default-deny attachment placement, and blackhole/segmentation tests. Preserve out-of-band management and rollback routes. A rollback that restores JSON but leaves propagated/service state changed is incomplete.
Cost and quota model
Model per account/VPC/Region/AZ: VPC/TGW/Cloud WAN attachments and processing, peering/inter-Region transfer, NAT hours/data, interface endpoint hours/AZ/data, Resolver endpoint IP hours/query, Network Firewall endpoint hours/GB, GWLB/appliance/licensing, Direct Connect ports/data/partner, VPN hours/data, public IPv4, IPAM, flow/DNS/firewall log ingestion/storage/query, load balancers, cross-AZ/Region and internet egress.
Trace each representative flow and charge at every hop. Centralization can add TGW + firewall + NAT + cross-AZ processing to one byte. Distributed components can cost more idle hours. Assign chargeback and anomaly alerts.
Track VPC/subnet addresses, routes, TGW tables/routes/attachments, Cloud WAN edges/segments/attachments/policy size, security rules/prefix lists, interface endpoints/policy size, NAT connections/ports, firewall capacity/throughput, Resolver endpoint IP/QPS/rules/associations, DX/VPN routes/BGP, flow-log delivery, and ENIs. Request capacity before migration.
Read-only evidence audit
With approved roles, redact account IDs, CIDRs, routes, endpoint IPs, domains, ASNs, firewall rules, and topology:
aws ec2 describe-vpcs --region ap-south-1
aws ec2 describe-transit-gateways --region ap-south-1
aws ec2 describe-transit-gateway-attachments --region ap-south-1
aws ec2 describe-transit-gateway-route-tables --region ap-south-1
aws ec2 describe-vpc-endpoints --region ap-south-1
aws ec2 describe-ipams --region ap-south-1
aws route53resolver list-resolver-endpoints --region ap-south-1
aws route53resolver list-resolver-rules --region ap-south-1
aws network-firewall list-firewalls --region ap-south-1
aws networkmanager list-core-networks
aws directconnect describe-connections
aws ec2 describe-vpn-connections --region ap-south-1
Lists do not prove routes or traffic. Follow pagination, inspect associations/propagations/policies/metrics/logs, and run approved path tests from both directions.
Troubleshooting matrix
| Symptom | Evidence order | Frequent cause |
|---|---|---|
| attachment available, no traffic | source route, TGW/segment association/propagation, destination/return routes | attachment mistaken for routing |
| firewall drops return flow | forward/return path, AZ endpoint, appliance mode, NAT order | asymmetry |
| spoke cannot resolve on-prem domain | VPC Resolver, rule association/match, endpoint/target health, routes/SG | DNS path incomplete |
| central endpoint resolves public IP | private DNS/PHZ/profile association and resolver path | managed PHZ scoped to endpoint VPC |
| new VPC overlaps acquisition | IPAM allocation/history, imported/discovered space | unmanaged allocation |
| DX failure does not use backup | BGP advertisements/preference, VPN health, route tables | failover never tested |
| Cloud WAN attachment isolated | tag/metadata, first-match rule, acceptance, policy live state | no attachment-policy match |
| TGW shows route but app times out | SG/NACL/firewall/MTU/host/app and return path | control-plane route overclaimed |
| network bill spikes | per-flow hop/AZ/Region trace and duplicate endpoints/logs | hidden processing/transfer |
| shared VPC deletion blocked | participant resources/ENIs and ownership inventory | lifecycle contract absent |
Failure and decommissioning
Game-day one AZ's NAT/firewall/Resolver IP, TGW route error, Cloud WAN policy/tag error, IPAM exhaustion, overlapping route, DX location, VPN tunnel, on-prem DNS, inspection overload, endpoint policy deny, Region isolation, log-delivery failure, and compromised network pipeline. Measure detection, path change, state/session loss, business impact, cost, recovery, and rollback.
Before detaching/unsharing/deleting, inventory participant resources, ENIs, routes, DNS associations, endpoints, firewall dependencies, hybrid advertisements, IP allocations, logs, and consumers. Migrate and prove alternate paths, freeze creation, remove propagation/routes/associations, verify denied/healthy paths, detach, release addresses after quarantine, and retain evidence. Never delete a shared VPC/TGW/rule because its owner view appears empty.
Practical lab
Download the AWS262 multi-account network architecture pack. It contains a 30-account/three-Region case, flow/route matrix, DNS/IPAM workbook, cost/quota model, ten failure cases, and answer directions.
Submit account/VPC ownership ADR; IPv4/IPv6 plan; TGW versus Cloud WAN decision; segment/route tables; RAM lifecycle; hybrid/DNS/endpoints; ingress/egress/inspection; security/evidence; cost/quotas; automation/rollback; all case diagnoses; and game-day/decommission proof.
Knowledge check
- When is shared VPC appropriate, and what remains shared risk?
- How do TGW association and propagation differ?
- Why must VPC route tables accompany TGW routes?
- When does Cloud WAN add value over Regional TGWs?
- Why are attachment tags security-sensitive in Cloud WAN?
- What does IPAM delegation/RAM sharing provide?
- Which IPv6 controls remain necessary?
- Why should workloads use VPC Resolver rather than outbound endpoint IPs?
- What breaks centralized interface-endpoint private DNS?
- Why can distributed endpoints be safer despite higher idle cost?
- What makes stateful inspection symmetric?
- Which independent failures must hybrid links tolerate?
- Why do route evidence and flow-log
ACCEPTnot prove application success? - How do central egress charges accumulate per hop?
- What sequence safely decommissions a shared network?
Lesson acceptance
Pass requires requirement/flow inventory; VPC ownership ADR; address/IPAM/IPv6 plan; connectivity comparison; explicit TGW/Cloud WAN segmentation and route behavior; DNS; hybrid; endpoints/PrivateLink; ingress/egress; symmetric inspection; layered security; observability; quotas/cost per flow; controlled automation/rollback; ten failure diagnoses; failover tests; and dependency-safe decommissioning.
Official sources
- Scalable, secure multi-VPC infrastructure
- VPC sharing
- How Transit Gateway works
- AWS Cloud WAN concepts
- Cloud WAN core-network policy
- Cloud WAN service insertion
- IPAM Organizations integration
- Route 53 Resolver rules
- Centralized VPC endpoints
- Centralized network inspection
- Avoid Network Firewall asymmetric routing
- Direct Connect resilience