Lesson 262 · AWS Learning Path

AWS 262: Multi-account network architecture

· Published · 14 min read

Labelled process diagram for AWS 262: Account and connectivity requirements to Shared or distributed network platform to Workload VPC paths to Segmentation, inspection, and operations evidence, with decision, proof...

Why this matters to an architect

An AWS account boundary does not create, prevent, or inspect network traffic. A secure multi-account design must deliberately join VPC ownership, addresses, routing domains, DNS, hybrid links, ingress/egress, private service access, inspection, identity, telemetry, capacity, cost, and recovery.

Centralizing everything can reduce duplicated appliances but creates shared blast radius, policy scale, inter-AZ/Region paths, and an operations bottleneck. Distributing everything can improve isolation and autonomy but multiplies endpoints, routes, cost, and configuration drift. The architect chooses per capability from measurable requirements - not from a fashionable hub diagram.

Outcomes

By the end, you can:

  • derive network requirements and trust zones from workload flows;
  • choose per-account VPCs, shared VPCs, or a portfolio hybrid;
  • compare peering, Transit Gateway (TGW), and Cloud WAN;
  • design route tables/segments for isolation, shared services, inspection, and hybrid paths;
  • plan IPv4/IPv6 with delegated VPC IPAM and RAM-shared pools;
  • design centralized or distributed DNS, endpoints, ingress, egress, and firewalls;
  • preserve stateful-routing symmetry and multi-AZ/Region failure domains;
  • design redundant Direct Connect/VPN and BGP policy;
  • assign owner/participant responsibilities and automate safe changes;
  • prove reachability, denied paths, DNS, performance, telemetry, cost, failover, and rollback.

Begin with flows and failure requirements

Inventory every required and forbidden path:

DimensionQuestions that change design
source/destinationaccount, VPC/subnet, on-premises/site, internet, partner, AWS service, SaaS
protocolIPv4/IPv6, TCP/UDP/ICMP, ports, DNS, multicast, MTU/fragmentation
trust/dataenvironment, tenant, data class, regulatory/residency, source identity
performancelatency, jitter, throughput, packets/flows, connections, DNS QPS
resilienceAZ/Region/site/device failures, RTO/RPO, degraded operation
inspectionstateful/stateless, TLS, IDS/IPS, source/destination preservation, fail-open/closed
ownershiprequester, approver, route/DNS/firewall owner, workload responder
evidenceflow/firewall/DNS logs, route/config history, active probes, packet evidence
costhourly attachments/endpoints/NAT/firewall, processing, inter-AZ/Region/egress

Draw the forward and return path hop by hop. Include DNS before connection establishment and identity/service policy before data flow. “Connected to TGW” is not an application-flow design.

Account and VPC ownership models

VPC per workload account

The workload account owns VPC/subnets/routes/security groups and application resources, while central IPAM, TGW/Cloud WAN, Resolver rules, firewall policy, and guardrails supply the platform. This gives strong ownership/failure boundaries and independent endpoint/security policy, but increases VPCs, attachments, address use, endpoint/NAT cost, and lifecycle automation.

Use when workloads need independent routing, sensitive isolation, delegated change cadence, or separate lifecycle. Standardize with account vending/IaC; do not allow arbitrary CIDRs and unmanaged internet gateways.

Shared VPC

The network account owns a VPC and shares selected subnets through RAM with participant accounts. Participants create supported application resources in those subnets, own their resources/security groups, and cannot manage owner/other-participant resources. The owner controls VPC CIDRs, subnets, route tables, NACLs, gateways, DHCP/DNS attributes, and VPC-level telemetry.

Benefits include fewer VPCs/attachments, efficient IPv4 use, simple same-VPC connectivity, and no inter-account data-transfer charge for same-AZ instance traffic. Tradeoffs include shared IP/routing/NACL/DNS failure domains, owner bottleneck, quota contention, reduced participant control, and unsupported resource types. Sharing separate subnets per participant improves NACL/routing separation but is not equivalent to a VPC boundary.

Use for cooperative teams with similar trust/connectivity/lifecycle requirements. Define IP allocation, AZs, route/NACL/security-group responsibility, flow-log access, quota, incident, participant removal, cost, and resource deletion order.

Hybrid portfolio

Common enterprise answer: dedicated VPCs for regulated/production/unique routing, shared VPCs for aligned platform or development portfolios, and TGW/Cloud WAN/PrivateLink between domains. Document why each portfolio shares a failure domain.

Choose the connectivity mechanism

MechanismGood fitReject or constrain when
VPC peeringsmall number of direct, non-transitive VPC relationshipsmesh scale, overlapping CIDRs, transitive hub/inspection needed
Transit GatewayRegional hub-and-spoke, multiple route domains, VPC/VPN/DX/Connect, centralized services/inspectionglobal policy consistency is dominant or per-flow narrow service exposure is enough
Cloud WANpolicy-driven multi-Region/site core, global segments, attachment policy, service insertion, centralized visibilitysmall single-Region estate where operational/cost complexity is unjustified
PrivateLinkexpose one service privately with producer/consumer isolation and no broad routinggeneral bidirectional network connectivity is required
VPC Latticeservice-network connectivity and policy for supported application servicesarbitrary network protocols/routing or unsupported resources
shared VPCparticipant resources need one owner-controlled VPCstrong VPC-level isolation/autonomy needed

Do not select solely by maximum scale. Compare operational model, segmentation, propagation control, route convergence, multi-Region behavior, appliance insertion, IPv6, quotas, observability, cost, migration, and rollback.

Transit Gateway route-domain design

A TGW is Regional. It routes IPv4/IPv6 between attachments using TGW route tables. Each attachment associates with one route table and can propagate routes to multiple route tables. VPC subnet route tables still require routes to TGW, and destination/return routes must exist.

Do not leave every attachment in default association/propagation. Design route tables such as:

Prod attachment -> Prod ingress table -> approved Prod/shared/inspection paths
NonProd attachment -> NonProd table -> isolated NonProd/shared paths
Inspection attachment -> Inspection table -> spoke/hybrid return routes
Hybrid attachment -> Hybrid table -> approved summarized AWS prefixes
Quarantine attachment -> no propagation; explicit remediation services only

Static routes override propagated routes for the same destination; blackhole routes intentionally drop matching traffic. TGW peering requires static routes. VPC propagation advertises VPC CIDRs; BGP attachments introduce learned routes. Record association, propagation, static/blackhole, prefix ownership, and expected route evaluation.

TGW sharing through RAM lets participant accounts create VPC attachments; either side can delete an attachment. Approval, tags, route association/propagation, change ownership, and deletion alarms must be automated. A participant-created attachment must not automatically join a privileged default table.

Cloud WAN design

Cloud WAN creates a global/core network with Regional core network edges defined by a versioned JSON policy. Segments are isolated routing domains by default. Attachment-policy rules map attachments using attachment tags/metadata, in rule-number order; first match wins, and a miss can leave the attachment unassociated. Acceptance can be required.

The policy defines Regions, segments, sharing, routing, attachment policies, network-function groups, and service insertion. Policy versions and change sets support review/rollback. Current policy version 2025.11 is required for newer routing policies/BGP-community capabilities; verify feature/Region/attachment support rather than copying syntax blindly.

Use RAM to share the core network. The core owner controls policy, edges, acceptance, and global routing; attachment owners manage their approved attachments/tags. Because tags can place an attachment in a segment, protect tag mutation and default-deny unmatched/malformed cases.

Cloud WAN does not remove the need for VPC routes, DNS, address planning, firewall behavior, endpoint policy, hybrid redundancy, or workload security controls.

IP addressing and VPC IPAM

Create a hierarchy from enterprise space to environment/Region/portfolio pools. Delegate an IPAM member account through the supported IPAM Organizations workflow so its service-linked role can monitor organization use; merely registering generic trusted access incorrectly can miss required setup. Share pools through RAM to accounts/OUs.

Define allocation locale, allowed netmask, auto-import/discovery, required tags, provisioned CIDRs, utilization thresholds, reservation/growth, BYOIP/public space, release quarantine, and M&A overlap. Account vending requests CIDRs idempotently and records allocation/resource/account/Region owner.

IPv4 plans must include endpoints, TGW/appliance/DNS subnets, scaling, blue-green, DR, acquisitions, and on-premises. Avoid giant VPCs “for future use” and tiny subnets that block scaling. NAT is not an address-management strategy.

Plan IPv6 intentionally: dual-stack versus IPv6-only subnets, egress-only internet gateway, DNS/AAAA, NAT64/DNS64 where needed, security controls, logs/tools, hybrid routing, load balancer/service support, and applications. IPv6 removes IPv4 exhaustion, not segmentation or egress governance.

DNS architecture

Every VPC normally uses Route 53 VPC Resolver at its VPC-provided address. Do not configure workloads to send ordinary DNS directly to outbound Resolver endpoint IPs.

Central hybrid pattern:

  • multi-AZ inbound endpoints let on-premises resolvers query approved Route 53 private namespaces;
  • multi-AZ outbound endpoints forward matching domains to on-premises/partner resolvers;
  • forwarding/system rules are shared through RAM and associated with spoke VPCs;
  • private hosted zones or Route 53 Profiles/Global Resolver capabilities are associated/shared according to current support;
  • DNS Firewall rule groups, query logs, and domain ownership add control/evidence.

Forwarding a rule does not require spoke-to-endpoint VPC routing because the service handles it, but target DNS servers must be reachable from outbound endpoints. Define conditional-forwarding precedence, most-specific matches, split horizon, delegation, overlapping zones, loops, NXDOMAIN, TTL/cache, DNSSEC, endpoint IP/capacity, and failover.

Unsharing/deleting a Resolver rule changes associated VPC behavior to remaining rules. Treat as a production dependency. Test AWS-to-on-premises, on-premises-to-AWS, denied domains, target/endpoints failure, and alternate path.

Private access to AWS and application services

Gateway endpoints provide private S3/DynamoDB routing with no endpoint hourly fee; route and endpoint/resource policies remain critical. Interface endpoints use PrivateLink ENIs, security groups, endpoint policy, per-AZ hourly and data-processing charges.

Distributed interface endpoints give each VPC independent availability/policy/blast radius but multiply cost and operations. Centralized endpoints reduce copies but add TGW/routing/DNS/data processing, policy-size/least-privilege complexity, and shared blast radius. Managed private DNS for an interface endpoint applies to its endpoint VPC; central consumers may require carefully managed private hosted-zone/profile/Resolver design. Compare total path cost and failure, not endpoint hourly price alone.

PrivateLink exposes a service without transitive routing or overlapping-CIDR concern and preserves a narrow trust boundary. It is often safer than connecting entire partner/tenant VPCs. Define endpoint-service permissions, acceptance, NLB/GWLB health, AZs, DNS, consumer endpoint policy/security groups, logging, quotas, and producer/consumer billing.

Internet ingress and egress

Ingress choices include CloudFront/WAF/Shield to regional ALB/API Gateway, Global Accelerator, public NLB/ALB, or controlled appliances. Prefer service-edge protection and application authentication over routing public traffic through a universal network hub. Define certificates/DNS, origin restriction, client IP, health, DDoS, IPv6, logs, failover, and cross-zone/Region cost.

For egress, compare distributed NAT gateway/firewall per VPC with centralized egress VPC through TGW/Cloud WAN. Centralized egress improves policy and public-IP control but creates shared capacity/failure, longer paths, TGW/firewall/NAT/inter-AZ processing, and source attribution concerns. Deploy NAT/firewall endpoints per used AZ and keep AZ-local routing where required; cross-AZ “HA” can increase cost and hide an undersized design.

Use gateway/interface endpoints to avoid internet/NAT where suitable. Domain filtering, proxying, Network Firewall, or third-party controls must align with encrypted DNS/TLS and application requirements. Decide fail-open/closed explicitly.

Centralized inspection and routing symmetry

AWS Network Firewall and stateful appliances require both directions of a flow to traverse the same firewall endpoint/state context. In a TGW inspection VPC, use dedicated TGW attachment and firewall subnets/route tables, enable TGW appliance mode on the inspection attachment, and route forward/return paths through the firewall. Network Firewall does not support asymmetric routing.

Inspect selected trust-zone crossings rather than every packet by default. Validate source/destination preservation, SNAT order, fragment/MTU, east-west/on-premises/internet flows, endpoint scaling, rule capacity, Suricata/stateless-stateful ordering, TLS limitations, logging, and bypass routes. A green firewall endpoint does not prove traffic traverses it.

Cloud WAN service insertion uses network-function groups and policy actions to redirect selected same/cross-segment paths; tags and policy become high-risk routing controls. Test the deployed policy and actual flow path.

Hybrid and multi-Region connectivity

Direct Connect needs redundant connections, routers, devices, and preferably locations according to the required resiliency model. Maximum resiliency uses separate connections terminating on separate devices in more than one location. VPN can provide backup or primary paths; test health and failover rather than assuming BGP handles everything.

Document Direct Connect gateway/TGW/core associations, virtual interfaces, ASNs, BGP communities/local preference/AS path/MED, allowed prefixes, route limits, summarization, default-route policy, MTU, MACsec where applicable, encryption requirements, and maintenance ownership. Use the Direct Connect failover test and active network probes.

TGW peering, Cloud WAN, inter-Region VPC peering, and service-specific replication solve different needs. Build Regional failure domains: do not send all regions' egress, DNS, inspection, or endpoints through one Region unless the business accepts that dependency and latency/data-transfer cost. Define route convergence, stateful-session loss, DNS failover, control-plane access, and degraded local operation.

Security layers and data perimeter

Use routing/segments to limit reachability, NACLs for stateless subnet controls where justified, security groups for stateful workload ENI policy, endpoint/resource policies for service access, Network Firewall/GWLB for inspection, WAF for HTTP, and IAM/SCP/RCP for API/resource authorization. Network reachability is not authorization; IAM authorization is not network reachability.

Security-group referencing has topology/Region/service constraints; verify current TGW/Cloud WAN and cross-account support. Prefix lists simplify approved CIDRs but require version/change governance. Data perimeter controls may use VPC endpoints and aws:SourceVpce/aws:SourceVpc or organization context where supported; test service-to-service exceptions and break glass.

Observability and evidence

Collect VPC/TGW flow logs, Network Firewall flow/alert logs, DNS query logs, ELB/CloudFront/WAF logs, Direct Connect/VPN/TGW/Cloud WAN metrics/events, IPAM compliance/utilization, Reachability Analyzer/Network Access Analyzer findings, CloudTrail changes, AWS Config, and active synthetic probes.

Flow logs are sampled/aggregated records, not packet captures, and ACCEPT does not prove application success. Correlate DNS, routes, security groups/NACLs, firewall/NAT, load balancer, host, and application evidence with timestamps. Protect central logs as in AWS260.

Safe automation and change

Keep IPAM, VPC/subnet, RAM shares, attachments, route tables, Cloud WAN policies, DNS, endpoints, firewall, hybrid configuration, and monitoring in versioned pipelines with one authoritative controller per resource. Validate schemas and route impact, model forward/return paths, run policy/reachability tests, deploy to a network test account/segment, canary attachments/routes, monitor, then batch.

Prevent route leaks with approved-prefix registries, summarization, max-prefix/route quotas, default-deny attachment placement, and blackhole/segmentation tests. Preserve out-of-band management and rollback routes. A rollback that restores JSON but leaves propagated/service state changed is incomplete.

Cost and quota model

Model per account/VPC/Region/AZ: VPC/TGW/Cloud WAN attachments and processing, peering/inter-Region transfer, NAT hours/data, interface endpoint hours/AZ/data, Resolver endpoint IP hours/query, Network Firewall endpoint hours/GB, GWLB/appliance/licensing, Direct Connect ports/data/partner, VPN hours/data, public IPv4, IPAM, flow/DNS/firewall log ingestion/storage/query, load balancers, cross-AZ/Region and internet egress.

Trace each representative flow and charge at every hop. Centralization can add TGW + firewall + NAT + cross-AZ processing to one byte. Distributed components can cost more idle hours. Assign chargeback and anomaly alerts.

Track VPC/subnet addresses, routes, TGW tables/routes/attachments, Cloud WAN edges/segments/attachments/policy size, security rules/prefix lists, interface endpoints/policy size, NAT connections/ports, firewall capacity/throughput, Resolver endpoint IP/QPS/rules/associations, DX/VPN routes/BGP, flow-log delivery, and ENIs. Request capacity before migration.

Read-only evidence audit

With approved roles, redact account IDs, CIDRs, routes, endpoint IPs, domains, ASNs, firewall rules, and topology:

aws ec2 describe-vpcs --region ap-south-1
aws ec2 describe-transit-gateways --region ap-south-1
aws ec2 describe-transit-gateway-attachments --region ap-south-1
aws ec2 describe-transit-gateway-route-tables --region ap-south-1
aws ec2 describe-vpc-endpoints --region ap-south-1
aws ec2 describe-ipams --region ap-south-1
aws route53resolver list-resolver-endpoints --region ap-south-1
aws route53resolver list-resolver-rules --region ap-south-1
aws network-firewall list-firewalls --region ap-south-1
aws networkmanager list-core-networks
aws directconnect describe-connections
aws ec2 describe-vpn-connections --region ap-south-1

Lists do not prove routes or traffic. Follow pagination, inspect associations/propagations/policies/metrics/logs, and run approved path tests from both directions.

Troubleshooting matrix

SymptomEvidence orderFrequent cause
attachment available, no trafficsource route, TGW/segment association/propagation, destination/return routesattachment mistaken for routing
firewall drops return flowforward/return path, AZ endpoint, appliance mode, NAT orderasymmetry
spoke cannot resolve on-prem domainVPC Resolver, rule association/match, endpoint/target health, routes/SGDNS path incomplete
central endpoint resolves public IPprivate DNS/PHZ/profile association and resolver pathmanaged PHZ scoped to endpoint VPC
new VPC overlaps acquisitionIPAM allocation/history, imported/discovered spaceunmanaged allocation
DX failure does not use backupBGP advertisements/preference, VPN health, route tablesfailover never tested
Cloud WAN attachment isolatedtag/metadata, first-match rule, acceptance, policy live stateno attachment-policy match
TGW shows route but app times outSG/NACL/firewall/MTU/host/app and return pathcontrol-plane route overclaimed
network bill spikesper-flow hop/AZ/Region trace and duplicate endpoints/logshidden processing/transfer
shared VPC deletion blockedparticipant resources/ENIs and ownership inventorylifecycle contract absent

Failure and decommissioning

Game-day one AZ's NAT/firewall/Resolver IP, TGW route error, Cloud WAN policy/tag error, IPAM exhaustion, overlapping route, DX location, VPN tunnel, on-prem DNS, inspection overload, endpoint policy deny, Region isolation, log-delivery failure, and compromised network pipeline. Measure detection, path change, state/session loss, business impact, cost, recovery, and rollback.

Before detaching/unsharing/deleting, inventory participant resources, ENIs, routes, DNS associations, endpoints, firewall dependencies, hybrid advertisements, IP allocations, logs, and consumers. Migrate and prove alternate paths, freeze creation, remove propagation/routes/associations, verify denied/healthy paths, detach, release addresses after quarantine, and retain evidence. Never delete a shared VPC/TGW/rule because its owner view appears empty.

Practical lab

Download the AWS262 multi-account network architecture pack. It contains a 30-account/three-Region case, flow/route matrix, DNS/IPAM workbook, cost/quota model, ten failure cases, and answer directions.

Submit account/VPC ownership ADR; IPv4/IPv6 plan; TGW versus Cloud WAN decision; segment/route tables; RAM lifecycle; hybrid/DNS/endpoints; ingress/egress/inspection; security/evidence; cost/quotas; automation/rollback; all case diagnoses; and game-day/decommission proof.

Knowledge check

  1. When is shared VPC appropriate, and what remains shared risk?
  2. How do TGW association and propagation differ?
  3. Why must VPC route tables accompany TGW routes?
  4. When does Cloud WAN add value over Regional TGWs?
  5. Why are attachment tags security-sensitive in Cloud WAN?
  6. What does IPAM delegation/RAM sharing provide?
  7. Which IPv6 controls remain necessary?
  8. Why should workloads use VPC Resolver rather than outbound endpoint IPs?
  9. What breaks centralized interface-endpoint private DNS?
  10. Why can distributed endpoints be safer despite higher idle cost?
  11. What makes stateful inspection symmetric?
  12. Which independent failures must hybrid links tolerate?
  13. Why do route evidence and flow-log ACCEPT not prove application success?
  14. How do central egress charges accumulate per hop?
  15. What sequence safely decommissions a shared network?

Lesson acceptance

Pass requires requirement/flow inventory; VPC ownership ADR; address/IPAM/IPv6 plan; connectivity comparison; explicit TGW/Cloud WAN segmentation and route behavior; DNS; hybrid; endpoints/PrivateLink; ingress/egress; symmetric inspection; layered security; observability; quotas/cost per flow; controlled automation/rollback; ten failure diagnoses; failover tests; and dependency-safe decommissioning.

Official sources

Advertisement