Lesson 277 · AWS Learning Path

AWS 277: Overlapping CIDR, route-domain, DNS, inspection, and egress design

· Published · 12 min read

Labelled process diagram for AWS 277: Overlapping estates and names to Isolation, translation, proxy, or service exposure to Inspection and DNS boundary to Controlled destination plus long-term address plan, with...

Why this lesson matters

Two networks can both work independently with 10.0.0.0/8, yet become ambiguous the moment they must communicate. A router cannot know whether 10.20.4.7 means an acquired-company database, an existing corporate service, or a local VPC host. Duplicate private DNS names add a second ambiguity even when translation makes packet routing possible.

There is no route-table setting that gives one destination address two meanings in the same route domain. Architects must renumber, keep domains isolated, translate one or both sides, proxy at the application layer, or expose a narrowly defined service. Each choice changes DNS, source identity, inspection, return routing, logging, scale, failure, and cost. This lesson develops a phased solution rather than presenting NAT as a permanent universal answer.

Outcomes

By the end, you can:

  • detect full, partial, aggregate, IPv4, IPv6, and DNS namespace overlap;
  • explain why TGW route tables and longest-prefix routing cannot disambiguate identical destinations;
  • select renumbering, isolation, NAT, proxy, PrivateLink, VPC Lattice, or application migration;
  • design translated source and destination spaces with deterministic bidirectional mappings;
  • preserve stateful inspection symmetry and useful identity through translation;
  • resolve duplicate private DNS names without forwarding loops or split-view surprises;
  • create separate internet, private, AWS-service, IPv4, and IPv6 egress controls;
  • govern new allocation with IPAM, Organizations, and infrastructure policy;
  • diagnose asymmetric, stale-DNS, port-exhaustion, and attribution faults; and
  • produce a migration plan that removes temporary translation debt.

Define the collision precisely

An overlap exists when one routed prefix intersects another. It can be:

CollisionExampleConsequence
Identicalboth use 10.0.0.0/16every destination is ambiguous
Partial10.0.0.0/16 and 10.0.8.0/21the more-specific route captures part of one estate
Aggregatecorporate 10.0.0.0/8, VPC 10.40.0.0/16broad corporate route conflicts with AWS child
Host/serviceduplicate 10.2.3.4/32static host route selects only one meaning
IPv6reused ULA or conflicting planned prefixdual-stack can reproduce IPv4 ambiguity
DNSboth own db.corp.internalsame name can return different or unreachable addresses

Inventory VPC primary/secondary CIDRs, subnets, Kubernetes pod/service ranges, TGW/Cloud WAN, VPN/DX advertisements, branches, partners, acquisitions, databases, appliances, client VPN pools, link-local use, IPv6, and DNS zones. Record owner, location, routability, lease/allocation source, utilization, dependencies, and retirement date.

An IPAM scope can intentionally contain disconnected networks that reuse space, but those scopes cannot later be merged without resolving overlap. “They never need to connect” is a business assumption requiring an owner and review date.

Why routing alone cannot solve it

Routers select a next hop from a destination prefix, usually longest-prefix match. They do not know which business meant 10.10.1.5. Separate TGW route tables can isolate consumers into distinct route domains, but any one source route table still cannot install two equal destination prefixes with different semantic meaning.

TGW does not route between VPC attachments with identical/overlapping CIDRs. If a newly attached VPC overlaps one already attached, routes for the new VPC are not propagated. A static route can select one attachment, but then all matching traffic selects that meaning. It does not create ambiguity-aware routing.

Cloud WAN segments likewise create separate domains; sharing overlapping attachment routes into one segment creates conflict rather than resolution. VPC peering rejects overlapping CIDRs. Separate VRFs/segments are useful containment, not communication.

Decision hierarchy

Prefer choices in this order:

  1. Renumber when ownership and time permit. This removes permanent ambiguity.
  2. Expose a service through PrivateLink, VPC Lattice, API, queue, or proxy when consumers need only a small capability.
  3. Keep route domains isolated when no communication is required.
  4. Translate deterministically when broad temporary connectivity is unavoidable.
  5. Use application gateways/proxies when protocol-aware controls and identity are more valuable than network transparency.
RequirementSuitable patternMain limitation
one TCP/API servicePrivateLink endpoint service/resource endpointnot general bidirectional routing
many HTTP/service workloadsVPC Lattice/API gateway/proxyapplication integration and policy
temporary broad accessNAT appliances/gateways with mapped rangesstate, scale, identity, complexity
no communicationseparate TGW/Cloud WAN segments/accountsduplicated shared services/operations
permanent enterprise integrationphased renumberingmigration effort and downtime risk

Do not translate an entire 10.0.0.0/8 if only one database port is needed. Narrow exposure reduces routes, attack surface, and migration debt.

Translation design

Assume Estate A and Estate B both use 10.20.0.0/16. Reserve non-overlapping transit aliases:

  • A appears to B as 100.64.20.0/24;
  • B appears to A as 100.64.40.0/24; and
  • only approved servers receive one-to-one mappings.

These are examples; verify CGNAT/shared ranges do not conflict with carriers, Kubernetes, appliances, or existing plans.

For each flow document original and translated tuples in both directions:

StageSourceDestination
A client startsreal A clientB alias
translation boundarytranslated A source if SNAT requiredreal B server after DNAT
B server repliesreal B servertranslated A identity
reverse translationB alias/real mapping restoredreal A client

DNAT changes destination meaning; SNAT guarantees the reply returns through the translator but hides the original source unless logs/application identity preserve it. One-to-one mappings improve attribution and inbound initiation but consume alias space. Many-to-one/PAT conserves addresses but adds port capacity and weakens identity.

Translation must be stateful and highly available. Deploy per AZ or use a supported appliance/GWLB pattern, preserve symmetric forward/return paths, size connections/ports/throughput, log pre/post tuples, test idle timeout and failover, and define fail-closed behavior. A NAT instance/appliance also needs source/destination check settings, scaling, patching, and route automation as applicable.

AWS NAT gateway is not arbitrary full-featured enterprise DNAT. Public NAT primarily provides outbound internet translation; private NAT can translate source addresses for private connectivity patterns, including overlapping networks when paired with deliberate routing. NAT gateways cannot be routed through VPC peering, and zonal placement matters. Confirm the exact supported path rather than assuming any NAT object solves bidirectional overlap.

Service exposure instead of routed integration

PrivateLink lets consumers connect to endpoint ENIs without routing provider CIDRs. It works with overlapping networks because the consumer addresses a local endpoint and the provider topology is hidden. Use NLB-backed endpoint services or current resource endpoints where suitable. Application authentication, DNS, TLS, consumer permission/acceptance, endpoint policy support, and provider health remain required.

VPC Lattice can expose multiple services/resources with service-network policy and discovery. An HTTP reverse proxy/API Gateway can terminate TLS, authenticate callers, rewrite hostnames, rate-limit, and produce application logs. Queues/events/object exchange remove synchronous network dependence.

These designs may not preserve arbitrary protocols or source IP and do not support provider-initiated general access. That limitation is often beneficial: it converts an uncontrolled network merger into a defined contract.

DNS with duplicate namespaces

Address translation without DNS design leaves users calling real overlapping addresses. Build a name registry with zone, authoritative owner, view, record type, real address, translated/endpoint address, TTL, and migration date.

Patterns:

  • distinct names such as service.estate-a.internal and service.estate-b.internal;
  • consumer-view private hosted zones returning translated aliases;
  • endpoint-specific PrivateLink/Lattice DNS names;
  • a migration alias that changes from translated to new real address after renumbering; and
  • Resolver forwarding rules only to a terminal authority for each suffix.

Never create loops where Estate A forwards corp.internal to B and B forwards it back. Resolver chooses the most-specific rule/zone. A matching private hosted zone with no requested record returns NXDOMAIN rather than public fallback. Cached positive and negative answers make cutover/rollback noninstantaneous.

Test exact FQDN and type against each Resolver endpoint, with UDP/TCP and cold cache. Correlate Resolver query logs, corporate DNS logs, translator logs, Flow Logs, and application evidence. DNS logs may show a real or alias address; only the translation log can prove the resulting tuple.

Inspection and identity

Place stateful inspection where both directions and useful addresses are visible. Decide whether policy matches pre-NAT identity, post-NAT destination, or both. A firewall after many-to-one SNAT cannot distinguish original hosts from packet headers. Preserve workload identity in mTLS, tokens, Proxy Protocol where applicable, or application gateway logs.

A typical translated path can be:

A workload -> A route domain -> pre-NAT inspection
  -> translation/GWLB appliance -> B route domain
  -> B inspection/security -> B workload

The return must traverse the same state tables in reverse. TGW appliance mode can keep an inspection attachment zonally symmetric, but it does not synchronize NAT appliance state or fix bypass routes. Keep translation, inspection, and NAT order explicit.

Policy should use the correct HOME_NET/address variables for both real and alias ranges. Threat intelligence and allowlists that see aliases need a mapping lookup. Retain mappings with timestamps, ports, flow identifiers, owner, and privacy controls for incident response.

Separate egress paths

Do not send every destination through the overlap translator.

DestinationPreferred path
approved overlapping private serviceendpoint/proxy or translation boundary
public internet IPv4zonal NAT plus inspection or approved proxy
internet IPv6egress-only IGW or approved proxy/inspection
AWS servicesgateway/interface endpoint with policy/private DNS
on-premises nonoverlapDX/VPN/TGW route domain
DNSapproved VPC Resolver/inbound-outbound endpoints

Use destination prefix lists, explicit routes, DNS policy, and endpoint controls. Prove no 0.0.0.0/0, ::/0, peering, secondary TGW, alternate DNS/DoH, public IP, or proxy bypasses inspection. Egress centralization can multiply TGW, firewall, NAT, cross-AZ, and logging charges.

NAT64/DNS64 helps IPv6-only clients reach IPv4 services; it does not resolve two identical IPv4 estates and introduces DNS/address-family behavior that must be tested.

IPAM prevention and governance

Use an organization-delegated IPAM account to monitor CIDRs across member accounts and share governed pools through RAM. IPAM shows managed, unmanaged, compliant, noncompliant, overlapping, and ignored resources. Ignoring a resource suppresses evaluation; it does not make overlap safe.

Create separate Regional/environment pools, allocation netmask rules, required tags, and infrastructure guardrails. Alarm on overlap/noncompliance and utilization metrics. Reserve acquisition, partner, pod, and transition ranges. IPAM retains address history that can help explain ownership, but historical data is not a substitute for current CMDB/application mapping.

Prevent manual VPC creation outside pools through approved provisioning, policy checks, Config/event controls, and exception workflow. Review secondary CIDRs and container ranges, not only primary VPC CIDRs.

Phased renumbering

  1. Inventory all addresses, names, dependencies, hard-coded values, certificates, ACLs, licenses, and logs.
  2. Allocate non-overlapping target CIDRs from IPAM.
  3. Add target VPC/subnets or supported secondary CIDR and build dual environment.
  4. Update infrastructure, DNS, service discovery, security, monitoring, backups, and clients.
  5. Introduce endpoint/proxy/translation compatibility for unmigrated consumers.
  6. Canary stateless services, then stateful tiers with replication/cutover plans.
  7. Lower DNS TTL ahead of controlled change where appropriate.
  8. Migrate clients and dependencies in waves, proving positive and negative flows.
  9. stop new use of old ranges and observe traffic until zero-use criteria pass;
  10. remove translation/routes/DNS aliases in dependency order, then release space under governance.

Renumbering is an application program, not only a network change. Hard-coded IPs can exist in databases, certificates, firewall objects, monitoring, backup agents, partner allowlists, license servers, scripts, and vendor appliances. Define rollback for each wave, including data reconciliation and DNS cache time.

Read-only evidence

export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
aws ec2 describe-vpcs +  --query 'Vpcs[].{Vpc:VpcId,Cidrs:CidrBlockAssociationSet[].CidrBlock}' +  --output json
aws ec2 describe-transit-gateway-route-tables --output json
aws route53resolver list-resolver-rules --output json
aws ec2 describe-vpc-endpoints --output json
aws ec2 describe-nat-gateways --output json
aws ec2 describe-ipams --output json

For an approved IPAM scope:

IPAM_SCOPE_ID="ipam-scope-approved-id"
aws ec2 get-ipam-resource-cidrs +  --ipam-scope-id "$IPAM_SCOPE_ID" +  --filters Name=overlap-status,Values=overlapping +  --output json

Also inspect TGW associations/propagations/routes, Cloud WAN segments, VPC subnet routes, NAT/appliance/GWLB state, endpoints, private hosted zones, Resolver rules/logs, firewall rules/logs, Flow Logs, IPAM history/compliance, CUR, and CloudTrail. Redact addresses and topology.

Diagnose from evidence

SymptomFirst checkLikely cause
new TGW attachment has no routeoverlap and propagationTGW suppressed overlapping VPC route
request reaches wrong estatesource route-domain tableequal/covering destination selected wrong meaning
request arrives, reply absentSNAT and reverse routesreply bypasses translator/state
one hostname works for some VPCszone/rule association and cacheambiguous split view or stale answer
logs show only translator addressNAT placement/mapping logsSNAT removed network source identity
failures under loadports/connections/throughputtranslation capacity exhausted
IPv4 controlled, IPv6 bypassesAAAA and ::/0 pathseparate family not governed
endpoint works but app deniesTLS/application identitynetwork solved, authorization did not
migration alias will not changeTTL/negative cache/authoritywrong resolver or cached record

Follow one 5-tuple through DNS, source route, route domain, inspection, pre/post translation, destination, and every reverse lookup. Never add broad static routes or disable inspection merely to make it pass.

Cost and temporary-debt register

Price TGW/Cloud WAN attachments and bytes, PrivateLink/Lattice endpoint hours/processing, NAT/GWLB/firewall hours and bytes, appliances/licenses, cross-AZ/Region transfer, Resolver endpoints/queries, DNS/log storage, duplicate migration environments, IPAM tier, and engineering/support.

For every temporary translation/proxy record an owner, business dependency, mapping range, capacity, monitoring, monthly cost, expiry, target non-overlapping design, and removal gate. A workaround without an exit criterion becomes permanent unowned infrastructure.

Conflict-remediation workbook

Repair a scenario with two 10.0.0.0/8 estates, duplicate corp.internal, asymmetric firewall paths, uncontrolled IPv4/IPv6 egress, and 30 required application flows.

Submit:

  • complete IPv4/IPv6/DNS collision inventory;
  • scored renumber/isolate/NAT/proxy/PrivateLink/Lattice decisions per service;
  • real/alias address registry and 30 bidirectional tuple traces;
  • route-domain and longest-prefix proof;
  • Resolver zone/rule/view and cache plan;
  • inspection/NAT order, identity, capacity and HA design;
  • separate internet/private/AWS-service/DNS egress matrix;
  • IPAM pool and prevention controls;
  • five-wave renumbering and rollback plan;
  • 12 failure/bypass tests, cost/debt register, RACI; and
  • before/after proof that no AWS resource changed.

Knowledge check

  1. Can separate TGW tables route two identical destination CIDRs from one source?

No. One route domain still needs one unambiguous next hop per destination.

  1. What happens when a newly attached TGW VPC overlaps an existing VPC?

Its overlapping routes are not propagated.

  1. Why is SNAT often used with overlap translation?

It forces replies through translator state, but hides original network source identity.

  1. Why does PrivateLink help?

Consumers reach local endpoint addresses without provider-CIDR routing.

  1. Does DNS64/NAT64 solve duplicate IPv4 estates?

No. It enables IPv6 clients to reach IPv4 destinations.

  1. What is the preferred long-term correction?

Governed unique addressing and phased renumbering.

Lesson acceptance

The lesson is complete only when the learner can:

  • inventory and classify every address and DNS collision;
  • prove why routing/segmentation alone cannot disambiguate identical destinations;
  • select renumbering, isolation, translation, proxy, or service exposure per flow;
  • calculate original/translated tuples and symmetric return paths;
  • design unambiguous DNS, identity-preserving inspection, egress, IPv6, and observability;
  • diagnose route, DNS, state, capacity, source-attribution, and bypass failures;
  • govern future allocations through IPAM and organization controls;
  • execute phased renumbering with compatibility, rollback, and removal gates;
  • complete the 30-flow remediation dossier and full cost/debt register; and
  • attest that no staging or production AWS resource changed.

Official sources

Advertisement