MX Records and Email Routing: DNS, Priority and Failover

· Published · 13 min read

SMTP sender performing DNS lookup and attempting MX 10 then MX 20 with temporary failure retry plus no MX address fallback and null MX outcomes

An MX record tells an SMTP sender which host accepts mail for a domain. It does not redirect a mailbox, test server health or guarantee instant failover. Reliable routing depends on authoritative DNS, address records, reachable SMTP services, recipient handling, retry behavior and monitoring. A correct design begins with protocol behavior and is verified externally before DNS changes.

Read an MX record correctly

example.com. 3600 IN MX 10 mx1.example.net.
example.com. 3600 IN MX 20 mx2.example.net.
mx1.example.net. 3600 IN A 192.0.2.10
mx1.example.net. 3600 IN AAAA 2001:db8::10

Lower preference numbers are tried first; they are not percentages or capacity weights. A final dot makes the target absolute in a zone file. The target must resolve to A/AAAA records. Avoid a CNAME as an MX target. Equal preference offers alternatives, but does not guarantee precise load distribution.

Follow the SMTP routing sequence

  1. Query MX for the recipient domain.
  2. Sort usable exchanges by preference and resolve addresses.
  3. Connect to TCP 25, negotiate SMTP and attempt the recipient.
  4. On temporary DNS, connection or SMTP failure, queue and retry, possibly using another host.
  5. On acceptance the receiver assumes responsibility; on permanent rejection stop retrying that recipient.

Remote sender retry schedules affect failover timing. A DNS change cannot force every queued MTA to reconnect instantly.

Understand no MX and null MX

StateMeaningAction
Valid MXUse listed exchangesResolve and attempt
No MXRFC 5321 implicit MXTry domain A/AAAA
MX 0 .RFC 7505 null MXDomain accepts no mail
Target lacks addressBroken routingTemporary failure and diagnosis

Null MX is explicit and different from accidentally forgetting MX.

Query multiple DNS views

dig +noall +answer MX example.com
dig +trace MX example.com
dig @1.1.1.1 MX example.com
dig A mx1.example.net
dig AAAA mx1.example.net
dig -x 192.0.2.10

Compare authoritative and public recursive answers. Check delegation, serial, DNSSEC validation where enabled and TTL remaining in caches. Split-horizon DNS can intentionally differ internally, but internet senders require a correct public view.

Test the server behind DNS

nc -vz mx1.example.net 25
openssl s_client -starttls smtp -connect mx1.example.net:25 -servername mx1.example.net -crlf

EHLO probe.example
MAIL FROM:<[email protected]>
RCPT TO:<[email protected]>
RSET
QUIT

Use an authorized test recipient and stop before DATA when only routing is tested. Validate banner, EHLO, STARTTLS certificate, accepted domain and response time. A DNS answer alone does not prove the mail service works.

Protect every secondary MX

A higher-preference backup must know valid recipients, reject invalid ones, enforce abuse controls and relay only to authorized destinations. Accepting every address and later bouncing creates backscatter. Define encryption, queue limits, lifetime, monitoring and the protected route to the primary system.

Many services instead use several fully capable front ends at equal preference. Whatever design is chosen, every published host needs equivalent recipient and security policy.

Migrate MX without losing mail

  1. Inventory domains, current MX, TTL and accepted-recipient sources.
  2. Build and test the new service before publication.
  3. Lower TTL one or two existing TTL periods beforehand when needed.
  4. Add the new route while the current service remains available.
  5. Verify public DNS and external deliveries.
  6. Monitor both through cache and sender retry windows.
  7. Remove the former route only after queues, aliases and logs are reconciled.

Run a controlled failover exercise

Send uniquely identified messages to monitored test recipients from independent external systems. Disable one node in an approved window, then observe connection failure, alternate-host use, delay and eventual delivery. Restore it and check duplicates and gaps.

Test invalid recipients too. Every active MX should reject them consistently. Capture DNS, SMTP replies, Message-ID, Received headers and timing. A successful TCP connection is not an end-to-end test.

Monitor DNS and SMTP separately

  • Authoritative availability and consistent MX answers.
  • Target A/AAAA resolution and DNSSEC state.
  • External TCP/25 reachability.
  • EHLO, STARTTLS and certificate behavior.
  • Known-recipient acceptance and invalid-recipient rejection.
  • Queue depth, oldest message, disk and delivery latency.
  • Unexpected MX, TTL or certificate changes.

A running daemon can still reject all valid recipients after its directory fails.

Secure the receiving edge without breaking SMTP

Restrict administrative interfaces separately from public SMTP, patch the MTA and TLS library, rate-limit abusive patterns and log recipient enumeration attempts. Port 25 must remain reachable for internet mail, but submission ports and management services need authenticated, limited access. Never expose an open relay.

STARTTLS is normally opportunistic unless enforced through a policy such as MTA-STS or DANE in an appropriate design. Monitor certificate name and validity, but remember that ordinary SMTP delivery can proceed without encryption when no stronger policy applies. Document the intended transport policy so a stricter node does not behave differently from its peers.

Diagnose common MX failures in order

SymptomCheck firstLikely boundary
NXDOMAINDelegation and queried domain spellingDNS name
MX answer but no connectionA/AAAA, firewall and TCP 25 routeNetwork/service
TLS warningCertificate name, chain, time and policyTransport security
550 unknown userRecipient directory and domain ownershipRecipient policy
451 repeatedQueue, directory, storage and downstream routeTemporary receiver condition

Capture the complete DNS response and SMTP exchange. A shortened error string often removes the hostname or enhanced status detail needed to find the responsible system.

Keep DNS and mail changes reversible

Record the current zone values, provider configuration and validation evidence before a change. Use peer review for preference, hostname and TTL edits. Set a rollback condition based on failed external delivery, not on an arbitrary elapsed time. A rollback also requires the former service to remain operational and its accepted-recipient data to stay current.

After the change, verify each authoritative nameserver directly and compare several public resolvers. Retain old server logs until delayed and queued deliveries have aged out. Close the change only after inbound volume, rejects, latency and security controls match the intended baseline.

MX routing checklist

  • Use absolute targets with valid A/AAAA.
  • Interpret lower preference as preferred, not weighted.
  • Understand implicit and null MX.
  • Verify authoritative and recursive answers.
  • Test port 25, EHLO, STARTTLS and recipients externally.
  • Prevent relay and backscatter on secondary hosts.
  • Keep old routes through cache and retry windows.
  • Exercise failover with evidence.
  • Monitor end-to-end acceptance and queues.

Trace MX resolution from delegation to usable addresses

dig +trace MX example.com
dig @ns1.example.net example.com MX +noall +answer
dig mx1.example.net A +short
dig mx1.example.net AAAA +short
dig -x 192.0.2.10 +short

An SMTP sender queries MX, orders usable exchanges by preference and resolves their A/AAAA records. Lower preference numbers are tried first; equal preference offers alternatives rather than guaranteed traffic percentages. MX targets must be hostnames with address records, not IP literals, and using a CNAME target is not a sound interoperable design.

Compare every authoritative server and public recursive views. A zone can be correct on one nameserver and stale on another. Record delegation, SOA serial, DNSSEC validation where used, TTL and exact answers.

Distinguish explicit MX, implicit MX and null MX

DNS stateSender behaviorOperator meaning
One or more usable MXAttempt exchanges by preferencePublished receiving service
No MX recordImplicitly try domain address under SMTP rulesMail may still be attempted
MX 0 .Do not attempt deliveryDomain explicitly accepts no mail
MX target has no usable addressTemporary DNS/routing failureBroken publication, not null MX
NXDOMAINDomain does not existPermanent naming failure

Do not publish null MX on a domain that receives mail through hidden application paths. Conversely, omitting MX does not explicitly declare no mail because implicit behavior exists.

Test DNS, TCP, SMTP and recipient policy separately

nc -vz mx1.example.net 25
openssl s_client -starttls smtp \
  -connect mx1.example.net:25 \
  -servername mx1.example.net -crlf

EHLO probe.example
MAIL FROM:<[email protected]>
RCPT TO:<[email protected]>
RSET
QUIT

Use an authorized recipient and stop before DATA when testing only recipient routing. Capture banner, EHLO capabilities, STARTTLS certificate, response timing and full SMTP replies. Run externally because an internal split-DNS or firewall view can hide the public failure.

A successful TCP connection proves only port reachability. A daemon can be up while directory lookup fails, every valid recipient receives 451, or the host accidentally relays. Test valid and invalid recipients under approved conditions.

Design failover around sender retry behavior

MX preference is not active health checking. A sender may try another equal or higher-numbered exchange after connection or temporary failure, but retry schedule, caching and queue state belong to the remote MTA. DNS changes cannot force all queued senders to reconnect immediately.

DesignRequirementRisk
Equal-preference front endsEquivalent recipient/security policyOne inconsistent node causes intermittent failures
Higher-numbered backupCurrent recipient data and protected onward relayBackscatter or open relay if it accepts everything
Cold disaster MXContinuously tested DNS/TLS/recipient pathFails precisely during emergency
DNS-only failoverTTL and cache overlap planningOld and new routes coexist longer than expected

Every published host needs queue, disk, certificate, abuse and recipient monitoring. A backup is production infrastructure even when rarely selected.

Prevent secondary MX backscatter and relay abuse

A secondary that accepts every recipient and later generates bounces becomes a backscatter source. Synchronize valid-recipient or routing data, reject invalid recipients during SMTP where possible, and authorize onward relay only to the protected destination. Set queue lifetime, disk limits, TLS policy and alerting.

secondary_mx_policy:
  accepted_domains: explicit inventory
  recipients: synchronized directory or safe verification
  relay_to: authenticated/restricted primary route
  public_relay: denied
  invalid_recipient: reject during RCPT
  queue_age_alert: defined by service objective

Test external relay attempts and recipient enumeration controls in an approved lab. Never copy a primary configuration that assumes an internal trusted network onto a public backup host.

Migrate receiving routes without losing or duplicating mail

  1. Inventory domains, aliases, valid-recipient sources, current MX/TTL and queued dependencies.
  2. Build the new service, authentication, TLS, monitoring and downstream delivery before DNS.
  3. Lower TTL sufficiently ahead when a shorter cache is required.
  4. Add/test the new path while the former route remains current.
  5. Verify authoritative/public DNS and deliver from independent external senders.
  6. Monitor both routes through cache and remote retry windows.
  7. Remove old MX only after volume, queues, aliases and logs reconcile.

Keep recipient-directory updates synchronized during overlap. Rollback requires the former service to remain operational; a DNS record alone cannot restore a decommissioned MTA.

Align MX changes with MTA-STS, DANE and certificates

An MTA-STS cached policy may authorize only listed MX patterns. Advertising a new emergency host before adding it to policy can cause enforced senders to defer. DANE deployments depend on DNSSEC-validated TLSA records and require their own safe transition. Certificate SANs must match actual MX hostnames, not only the recipient domain.

ChangePrecondition
New MTA-STS MXAdd policy pattern, release policy ID and allow cache overlap
Certificate rotationDeploy complete trusted chain on every node
DANE key/cert changePublish TLSA overlap under valid DNSSEC
Remove hostDrain queues and account for cached MX/policy

Probe each address with correct SNI and compare TLS-RPT reports by MX host after changes.

Monitor the complete receiving path

  • Delegation, authoritative consistency, SOA serial and DNSSEC.
  • MX answers, target A/AAAA and unexpected changes.
  • External TCP/25, banner, EHLO and STARTTLS certificate.
  • Known-recipient acceptance and invalid-recipient rejection.
  • Queue depth, oldest message, disk and downstream delivery latency.
  • Directory freshness, relay restrictions and abuse anomalies.
  • MTA-STS policy endpoint and TLS-RPT failure categories.

Use end-to-end synthetic messages with unique Message-ID from independent systems, then verify receipt and headers. A green DNS probe does not prove delivery, and a received synthetic message does not prove every node is healthy.

Worked failure: backup MX accepts mail but cannot deliver it onward

The primary MX becomes unreachable during maintenance. Remote senders select the higher-numbered backup and receive 250. Operators assume failover succeeded, but the backup’s relay route uses an expired internal credential and its queue grows for hours. External monitoring checked only TCP/25.

The team preserves queue IDs and ages, restores the authenticated onward route and controls queue release to prevent a delivery burst. It verifies recipient suppression and duplicate behavior, then sends synthetic messages through each published MX. No mail is deleted merely to make queue graphs green.

Permanent controls add oldest-queue and downstream-delivery objectives, certificate/credential expiry monitoring, valid-recipient synchronization and a scheduled failover exercise. MX resilience is accepted only when a message travels from external sender through the alternate host into the final mailbox.

Keep a reversible DNS and mail change record

Record pre-change zone answers, service configuration, recipient-source version, certificate fingerprints, MTA-STS/DANE artifacts, external tests, owner and rollback condition. Verify each authoritative nameserver directly after publication and retain old logs until delayed mail ages out.

Close the change when inbound volume, invalid-recipient behavior, queue latency, downstream delivery and security controls match the intended baseline across every advertised host. A short successful test immediately after DNS edit is not sufficient evidence.

Test IPv4 and IPv6 paths independently

If an MX hostname publishes both A and AAAA, remote senders may choose either address according to their implementation and network state. A working IPv4 path does not excuse an unreachable IPv6 listener, wrong IPv6 firewall, missing reverse DNS or different SMTP configuration. Broken IPv6 can create intermittent delay that is difficult to reproduce from an IPv4-only monitor.

dig +short A mx1.example.net
dig +short AAAA mx1.example.net
nc -4 -vz mx1.example.net 25
nc -6 -vz mx1.example.net 25
openssl s_client -4 -starttls smtp -connect mx1.example.net:25
openssl s_client -6 -starttls smtp -connect mx1.example.net:25

Run probes from networks with real connectivity and correct hostname validation. Compare banner, EHLO, STARTTLS, certificate and recipient response across every address. If IPv6 is not production-ready, do not publish an AAAA record merely because the host has an address.

Distinguish DNSSEC validation failure from an unsigned answer

A resolver that validates DNSSEC treats a bogus chain differently from an unsigned domain. Expired signatures, broken delegation signer records, incorrect key rotation or inconsistent authoritative zones can make MX resolution fail for validating senders while a nonvalidating local lookup appears normal.

ObservationCheck
SERVFAIL at validating resolverUse dig +dnssec, trace delegation and inspect DS/DNSKEY/RRSIG
One nameserver differsSOA serial, zone transfer/deployment and signature set
Works with validation disabledDo not bypass; repair the DNSSEC chain
Intermittent after key rotationTTL overlap and publication sequence

Monitor signature expiry and validate changes before deployment. DNSSEC rollback must follow a safe provider-specific sequence; deleting records randomly can prolong a bogus state through caches.

Troubleshoot MX delivery in protocol order

  1. Confirm the recipient domain spelling, existence and delegation.
  2. Query MX at every authoritative server and compare recursive answers.
  3. Resolve every target A/AAAA and check routing/firewall to TCP 25.
  4. Capture banner and EHLO; verify expected hostname and capabilities.
  5. Negotiate STARTTLS with the MX hostname and inspect chain/SAN.
  6. Test a controlled valid and invalid recipient.
  7. Inspect receiving queues, directory and downstream route.
  8. Trace a uniquely identified message to final mailbox.
SymptomEarliest likely boundary
NXDOMAINName/delegation
MX answer, no addressTarget DNS
Connection refused/timed outNetwork, firewall or daemon
STARTTLS mismatchCertificate/SNI or wrong node
Valid user rejectedRecipient directory or accepted-domain policy
Accepted but not deliveredQueue/downstream mailbox

Stop at the earliest failing boundary. Changing MX preference cannot repair a stale recipient directory.

Design receiving capacity for normal and failover states

Measure connections, recipients, bytes, TLS handshakes, content-scanning cost, directory latency and downstream throughput by minute and node. Capacity must handle loss of a node without uncontrolled 421/451 responses or disk exhaustion. Use temporary SMTP failures when the system cannot safely accept responsibility; never return 250 and silently discard.

Set queue age and disk thresholds, reserve headroom for retries, and test backpressure. Remote senders can retry in bursts after recovery. Release queued mail in controlled order while preserving per-message expiry and avoiding duplicates. A load balancer health check should exercise SMTP readiness, not only TCP accept.

Record maintenance mode so a draining node stops receiving new sessions but completes or safely returns current transactions. Reconcile accepted recipients with queued/delivered outcomes during every failover exercise.

Assign ownership across DNS, network, MTA and directory

MX reliability crosses teams. DNS owns delegation and records; network owns routes and port 25; messaging owns SMTP, queues and relay; identity teams own accepted recipients; security owns TLS and abuse controls. Keep one service inventory and incident contact path. Test escalation during failover so operators can repair the earliest boundary without waiting to discover who controls it.

Review third-party receiving providers for domain verification, data retention, routing exports and exit procedure. Before contract termination, preserve aliases, recipient states, logs and overlap routing. Removing a vendor MX before the replacement accepts every required domain creates permanent bounces that DNS rollback may not immediately cure.

Final operational test

Verify every advertised address and complete one end-to-end delivery through each intended failover route before approval. Archive the evidence with the DNS release.

Primary references

Continue learning

Related technical notes

Technical review

Need this checked against your own sending system?

Share the domain, headers, bounces, provider warning, logs, or infrastructure symptom and NitWings will identify the practical next step.

Schedule a Technical Review
Advertisement