Lesson 272 · AWS Learning Path

AWS 272: Hybrid DNS with Route 53 Resolver

· Published · 15 min read

Labelled process diagram for AWS 272: Client and search suffix to Authoritative or forwarding decision to Inbound or outbound Resolver endpoint to Answer, cache, log, and failure evidence, with decision, proof and...

Why this lesson matters

An application can have healthy instances, correct routes, open security groups, and valid certificates yet remain unavailable because its name resolves incorrectly. Hybrid DNS joins AWS naming with corporate DNS across VPN or Direct Connect. It must answer two different questions: how clients outside AWS resolve private AWS names, and how workloads inside a VPC resolve names hosted by corporate DNS.

Amazon Route 53 VPC Resolver is the recursive DNS service available in every VPC. Resolver endpoints make that service reachable across a hybrid network. Rules decide which suffixes leave AWS. Private hosted zones, caches, endpoint security groups, routes, DNSSEC, and central governance then influence the answer. This lesson starts with Linux DNS concepts and develops an architect-level method for tracing and proving the complete query path without creating chargeable endpoints.

Outcomes

By the end, you can:

  • distinguish a stub resolver, recursive resolver, forwarder, and authoritative name server;
  • explain zones, delegation, TTL, negative caching, UDP, TCP, EDNS, and DNS over HTTPS;
  • use default and delegation inbound endpoints correctly;
  • use outbound endpoints with FORWARD, SYSTEM, RECURSIVE, and DELEGATE rule behavior;
  • calculate rule and private hosted zone precedence;
  • design endpoint IPs, Availability Zones, routes, security groups, and capacity;
  • centralize DNS resources with Route 53 Profiles and AWS Resource Access Manager;
  • use query logs, CloudWatch metrics, Flow Logs, and server logs as separate evidence layers;
  • diagnose NXDOMAIN, SERVFAIL, REFUSED, timeout, stale, and intermittent results;
  • model DNSSEC, split-view DNS, multi-account, multi-Region, failure, and cost boundaries; and
  • complete a no-change hybrid DNS architecture workbook suitable for review.

DNS foundations for a Linux operator

An application normally asks the operating system's stub resolver to resolve a name. On Linux, /etc/resolv.conf, NetworkManager, systemd-resolved, DHCP options, containers, and application-specific settings can influence which recursive server receives the query. A recursive resolver follows referrals or uses a configured forwarder, caches the result, and returns a final answer. An authoritative server owns a zone and returns records from that zone. A forwarder sends selected queries to another recursive resolver instead of walking the DNS hierarchy itself.

A zone is an administrative portion of the DNS namespace. Records include A for IPv4, AAAA for IPv6, CNAME for an alias, MX for mail routing, TXT for text, PTR for reverse lookup, SOA for zone authority metadata, and NS for delegated name servers. Delegation is not forwarding: delegation publishes authority with NS records, while conditional forwarding is resolver policy.

Every answer has a time to live, or TTL. Recursive resolvers may reuse a positive answer until its TTL expires. They can also cache a negative answer, usually using values derived from the zone's SOA record. A changed record may therefore appear correct to one client and stale to another. Search suffixes can silently transform db01 into several fully qualified names. Test with a trailing dot, such as db01.corp.example., when you need to exclude suffix expansion.

Classic DNS usually starts with UDP port 53. A large or truncated response can retry over TCP port 53; zone transfers use TCP. EDNS extends DNS message capabilities, including larger UDP payloads. Firewalls that allow UDP 53 but block TCP 53 produce failures that depend on response size. DoH carries DNS through HTTPS, normally TCP 443, and changes encryption and network inspection expectations.

cat /etc/resolv.conf
resolvectl status 2>/dev/null || true
dig app.aws.corp.example. A
dig @10.20.10.53 app.aws.corp.example. A +tcp
dig @10.20.10.53 missing.aws.corp.example. A +noall +answer +authority +comments

dig +trace follows public delegation and is not proof of a private forwarding path. Query the exact resolver IP involved, record UDP or TCP, and inspect response code, flags, authority section, answer, TTL, and query time.

The VPC Resolver baseline

VPC Resolver answers VPC-local names, records in associated private hosted zones, and public names through recursive resolution. A workload normally reaches it at the VPC network address plus two. For example, a 10.40.0.0/16 VPC uses 10.40.0.2. VPC attributes and DHCP configuration affect workload use of this resolver.

The plus-two address is a VPC-local access mechanism, not an on-premises DNS destination. Routing 10.40.0.2 across Transit Gateway, peering, VPN, or Direct Connect is not the supported hybrid design. Use inbound endpoint IP addresses for queries entering AWS.

On-premises client
  -> corporate recursive resolver
  -> conditional forward or NS delegation
  -> inbound endpoint IPs in two AZs
  -> VPC Resolver
  -> private hosted zone, VPC name, or recursive answer

VPC workload
  -> VPC Resolver at VPC+2
  -> most-specific associated Resolver rule
  -> outbound endpoint IPs in two AZs
  -> corporate DNS servers
  -> authoritative corporate zone

Endpoints are Regional resources made from elastic network interfaces in selected subnets. The control plane creates endpoints, rules, associations, and logs. The data plane receives and forwards queries. A successful API status proves configuration exists; it does not prove routes, firewalls, authoritative data, cache state, or the client result.

Inbound resolution from the corporate network

A default inbound endpoint accepts queries that corporate resolvers conditionally forward. For example, corporate DNS forwards aws.corp.example to two inbound endpoint IPs. Resolver then answers from an associated private hosted zone or its normal resolution behavior.

A delegation inbound endpoint supports a different ownership model. Corporate DNS delegates a subdomain to VPC Resolver using NS records and the inbound endpoint IPs as glue. Authority is delegated instead of merely forwarded. Delegation inbound endpoints currently use Do53 only, so do not promise DoH for this category.

For either pattern:

  • allocate endpoint IPs in at least two Availability Zones;
  • make each endpoint subnet reachable from every approved corporate resolver;
  • permit required source CIDRs and protocols in the endpoint security group;
  • permit return traffic through NACLs, firewalls, TGW, VPN, or DX;
  • give corporate DNS at least two destination IPs and test each directly;
  • associate private hosted zones with the VPCs from which Resolver must see them; and
  • document whether the parent uses conditional forwarding or NS delegation.

Default inbound endpoints can support Do53, DoH, DoH-FIPS, or documented combinations. Do53 uses port 53 without DoH's application-layer encryption. DoH and DoH-FIPS use encrypted HTTPS sessions. Confirm client compatibility, certificate behavior, port 443 policy, logging limitations, and migration sequencing before changing protocol. AWS documents a source-IP query-logging limitation for DoH/DoH-FIPS inbound endpoints, which matters for attribution.

Outbound resolution to corporate DNS

An outbound endpoint is used when a VPC workload asks VPC Resolver for a suffix that corporate DNS owns. A forwarding rule names the suffix, outbound endpoint, and corporate target resolver IP addresses and ports. Associate the rule with every VPC that should use it, directly, through sharing, or through a Profile.

The endpoint ENIs originate queries toward those targets. The path needs endpoint subnet routes to corporate servers, return routes to endpoint subnet CIDRs, egress security-group rules, stateful return allowance, and NACL/firewall policy. A workload security group does not directly open the endpoint-to-corporate flow.

Never configure corporate DNS to send the same suffix back to the same AWS inbound path:

VPC Resolver: corp.example -> corporate DNS
corporate DNS: corp.example -> AWS inbound endpoint
AWS Resolver: corp.example -> corporate DNS

This forwarding loop causes delay, SERVFAIL, high query volume, and capacity consumption. Draw an owner and terminal authoritative source for every suffix before implementation.

Rule types and precedence

Rule typeCreatorPurpose
FORWARDCustomerSend a matching suffix to target resolver IPs through an outbound endpoint
SYSTEMCustomer/serviceMake Resolver handle a subdomain that would inherit a broader forward
RECURSIVEVPC ResolverDefault Internet Resolver behavior for otherwise unmatched names
DELEGATECustomerReach delegated authoritative servers through an outbound endpoint after matching NS delegation

Resolver uses the most-specific matching domain. A rule for engineering.corp.example is selected over corp.example. A dot rule, ., forwards the broad remainder, but autodefined rules retain selected AWS and local namespaces unless deliberately overridden. Broad forwarding can break AWS functionality.

A SYSTEM rule is an exception to a broader forward. If corp.example goes to corporate DNS but cloud.corp.example should stay with VPC Resolver, use the more-specific system rule. RECURSIVE is service-created; customers cannot create an arbitrary recursive rule.

DELEGATE preserves delegation semantics. It uses an outbound endpoint when matching NS information is encountered, such as delegating a child of a private hosted zone to corporate authoritative servers. It is not configured with the same target IP list as a FORWARD rule; authoritative server names and glue/delegation drive resolution.

Private hosted zones also use most-specific namespace matching. If a matching private hosted zone exists but the requested record/type does not, Resolver returns NXDOMAIN rather than falling back to a public zone. A Resolver rule for the same domain takes precedence over the private hosted zone.

High availability, security, and capacity

Use endpoint IPs in at least two AZs. Two IPs are independent destinations, not an active/passive pair managed by DNS health checks. The calling resolver must retry another destination when one fails. Test maintenance, AZ route loss, VPN/DX loss, corporate server loss, and slow responses.

For Do53 inbound endpoints, allow inbound UDP and TCP 53 from approved resolver CIDRs. For outbound endpoints, allow egress UDP and TCP to corporate DNS target ports; responses return to endpoint source ports. AWS publishes broader security-group patterns for maximum throughput because connection tracking can reduce capacity. Start with least privilege, measure, and understand throughput before widening rules. DoH needs its HTTPS path and protocol controls.

Current default quotas include four endpoints per account per Region, six IPs per endpoint, six IPs per rule, 1,000 rules per account per Region, and 2,000 rule-to-VPC associations per account per Region; several are adjustable. Each endpoint IP can process up to 10,000 UDP queries per second under favorable conditions. Actual capacity depends on query size, protocol, target latency, response time, and connection tracking. Restrictive tracked flows or a Network Load Balancer path can reduce inbound capacity substantially.

Monitor per-IP query-volume metrics rather than dividing account-wide volume by IP count. AWS recommends adding interfaces before an endpoint interface exceeds 50 percent of capacity. Outbound Resolver sends redundant queries for availability, so endpoint egress counts need not equal application queries. TCP-heavy responses, retries, loops, slow targets, and high-cardinality clients alter sizing.

Multi-account governance with Route 53 Profiles

Route 53 Profiles apply common DNS configuration across VPCs and accounts in the same Region. A Profile can carry private hosted zones, forwarding and system rules, DNS Firewall rule groups, query-log configurations, and interface VPC endpoints. It can govern DNSSEC validation, reverse DNS behavior, and DNS Firewall failure mode. Profiles can be shared through AWS RAM.

Only one Profile can be associated with a VPC. Local configuration can coexist during migration. For the same domain, a local VPC resource wins over an equally specific Profile resource; otherwise the more-specific domain wins. Centralization reduces drift but increases blast radius because a Profile update propagates to all associated VPCs.

RoleResponsibility
Network accountEndpoint VPCs, hybrid routes, security groups, IP capacity
DNS platformRules, Profiles, zones, query logs, DNS Firewall
Application accountVPC approval, client settings, application tests
Corporate DNS teamForwarders, delegations, target servers, server logs
Security teamDNSSEC, filtering, log access, retention and privacy

Repeat Regional resources for multi-Region designs. Validate each Region's zone associations, target reachability, sharing, logging, and failover. A Region outage plan should not require manual changes on every client.

DNSSEC, filtering, and evidence

DNSSEC validation verifies trust chains and signatures for signed data. It does not encrypt queries; DoH addresses transport confidentiality. Enable VPC validation directly or through a Profile only after testing public and forwarded namespaces. A broken trust chain commonly becomes SERVFAIL, which can be mistaken for routing failure.

Zone signing, registrar DS records, validating recursive resolvers, and corporate forwarded zones are separate responsibilities. Private zones and internal delegation need an explicit trust model. Test signed-valid, signed-bogus, unsigned, and forwarded names.

Resolver DNS Firewall filters domain names processed by VPC Resolver. It does not block direct IP connections. Fail closed returns SERVFAIL when filtering cannot be evaluated; fail open favors availability. Record and test this decision.

Query logging can capture VPC-originated queries, inbound endpoint queries, outbound recursive activity, and DNS Firewall actions. Destinations are CloudWatch Logs, S3, or Firehose. Records may include VPC, source, name/type, response code/data, time, and firewall action.

Caching changes evidence: Resolver logs unique queries but not every repeat served from cache. Absence of a second log record does not prove the client stopped querying. Logs do not prove the client received or accepted a response. Correlate:

  • client dig or application evidence;
  • Resolver query logs;
  • endpoint CloudWatch metrics;
  • VPC/TGW Flow Logs;
  • corporate resolver logs;
  • zone SOA, NS, and record data; and
  • CloudTrail configuration changes.

DNS names can expose customers and internal systems. Apply least-privilege access, encryption, retention, redaction/export controls, and a privacy owner.

Read-only Console and CLI evidence

In the Route 53 console, choose the correct Region and inspect VPC Resolver, endpoints, rules, query logging, Profiles, private hosted zones, DNSSEC, and DNS Firewall. Record category/protocol, IP/AZ/subnet, VPC, security groups, status, suffix/type/targets, associations, ownership/share status, and timestamps. Do not create, edit, associate, or delete.

export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
aws route53resolver list-resolver-endpoints --output json
aws route53resolver list-resolver-rules --output json
aws route53resolver list-resolver-rule-associations --output json
aws route53resolver list-resolver-query-log-configs --output json
aws route53profiles list-profiles --output json
ENDPOINT_ID="rslvr-in-or-out-approved-id"
RULE_ID="rslvr-rr-approved-id"
aws route53resolver get-resolver-endpoint +  --resolver-endpoint-id "$ENDPOINT_ID" --output json
aws route53resolver list-resolver-endpoint-ip-addresses +  --resolver-endpoint-id "$ENDPOINT_ID" --output json
aws route53resolver get-resolver-rule +  --resolver-rule-id "$RULE_ID" --output json
aws route53resolver list-resolver-rule-associations +  --filters Name=ResolverRuleId,Values="$RULE_ID" --output json

Also inspect VPC DNS attributes, zone associations, endpoint ENIs/subnet routes/security groups, RAM shares, Profile associations, log destinations, metrics, and quota values. Redact IDs, internal domains, IPs, and topology before sharing.

Diagnose from the response code

SymptomFirst meaning to testTypical causes
NXDOMAINName does not exist in selected namespacemissing record, wrong suffix, private-zone capture, negative cache
NOERROR without answerName exists but requested type may notwrong A/AAAA expectation, empty nonterminal
SERVFAILResolver could not complete safelyDNSSEC, loop, failed authority, Firewall fail closed
REFUSEDServer received but rejected queryrecursion ACL, wrong source/server role
TimeoutNo usable response arrivedroute, SG/NACL/firewall, TCP/UDP, dead endpoint
Works in one VPCAssociation differsmissing rule/zone/Profile association, local conflict
Small answers onlyTransport defectTCP 53 blocked, EDNS/MTU issue
Alternating resultOne destination is unhealthyendpoint IP/AZ/target path defect
Old addressCache or split viewTTL, local cache, wrong authority tested

Use this order:

  1. Capture exact FQDN, type, client, time, resolver, protocol, and result.
  2. Query each configured resolver directly with UDP and TCP, avoiding a warm application cache.
  3. Determine the most-specific rule, local/Profile precedence, and private-zone match.
  4. For inbound, trace corporate DNS to every endpoint IP and confirm zone visibility.
  5. For outbound, trace every endpoint ENI to every target and return path.
  6. Correlate query logs, metrics, Flow Logs, target logs, and authoritative data.
  7. Check TTL/negative cache, DNSSEC, Firewall, delegation/glue, and split view.
  8. Change only the first proven failed control, then repeat positive and negative tests.

Never add a dot rule, allow DNS from everywhere, disable DNSSEC globally, or flush every production cache simply to make a test pass.

Hybrid DNS architecture workbook

Design this no-create scenario: two AWS Regions and twelve VPCs in three accounts use private names under aws.corp.example; corporate sites use corp.example; production may resolve corporate services, development may resolve only approved subdomains, and both directions must survive one endpoint IP/AZ failure.

Use at least 15 names covering valid A/AAAA, missing record, public/private overlap, PTR, delegated child, DNSSEC-valid and bogus, large TCP response, blocked domain, and broad/specific rule matches.

Evidence fieldRequired entry
Client and intentsource, FQDN, type, expected allow/deny
Stub configurationresolver, suffix behavior, cache state
Selected namespacerule/zone, local/Profile source, precedence
Directioninbound default/delegation, outbound forward/delegate, recursive
Network pathendpoint IP/AZ, route, SG/NACL/firewall, return
Authorityowner, zone, NS/glue or target resolver
SecurityDo53/DoH, DNSSEC, Firewall action
Evidenceclient, query log, metric, Flow Log, server log
Failure behaviorcode, retry/failover, TTL impact
Cost and ownerendpoint, query, logs, network services

Submit a suffix registry, bidirectional flow diagram, IP/AZ capacity sheet, rule/zone/Profile matrix, route/security matrix, observability plan, RACI, 15 rows, failure plan, rollback, and proof that no resource changed.

Cost, change safety, and rollback

Endpoint pricing is driven principally by endpoint IP address hours and queries processed, with Region-specific rates. Add Logs ingestion/retention, S3 or Firehose, DNS Firewall, traffic inspection, TGW/VPN/DX, applicable transfer, and tooling. More IPs improve capacity and resilience but add recurring cost. Loops can create an outage and an unexpected bill.

Before change, export endpoints, IPs, rules, associations, Profiles, zones, security groups, routes, logs, metrics, TTLs, and records. Use infrastructure as code, review, conflict simulation, canary VPCs, low-TTL preparation where appropriate, and success/error gates. Change one suffix path at a time.

Rollback restores the prior association, target set, Profile resource, NS/glue, policy, and record. It is not instantaneous because positive and negative caches can retain answers. Record the longest relevant TTL and recovery evidence.

This lesson creates nothing. Before/after inventory must show identical endpoint, IP, rule, association, Profile, zone, log, security-group, and route state.

Knowledge check

  1. Why should corporate DNS not query VPC plus-two directly?

It is VPC-local; hybrid ingress uses reachable inbound endpoint IPs.

  1. Forwarding versus delegation?

Forwarding is resolver policy; delegation publishes child authority with NS information.

  1. Which wins for api.dev.corp.example: corp.example or dev.corp.example?

The more-specific dev.corp.example rule.

  1. Why can a private name return NXDOMAIN instead of a public answer?

A matching private zone owns the namespace and Resolver does not fall back when its record is absent.

  1. Why are there fewer log events than client queries?

Repeats answered from Resolver cache are not logged as new unique queries.

  1. Why test UDP and TCP 53?

Large/truncated DNS answers retry over TCP.

  1. DNSSEC versus DoH?

DNSSEC validates data authenticity; DoH encrypts transport.

Lesson acceptance

The lesson is complete only when the learner can:

  • explain the Linux-to-authority DNS chain;
  • draw inbound default, inbound delegation, outbound forwarding, and outbound delegation;
  • predict rule, private-zone, local, and Profile precedence for 15 names;
  • prove two-AZ routing, security, retry, TCP/UDP, protocol, and capacity;
  • avoid loops and name an authoritative owner for each suffix;
  • diagnose response codes from client, Resolver, network, and server evidence;
  • design DNSSEC, Firewall, logging, privacy, multi-account, and multi-Region controls;
  • calculate endpoint, query, telemetry, and network cost ownership;
  • produce change and cache-aware rollback plans; and
  • attest that no staging or production AWS resource changed.

Official sources

Advertisement