MTA-STS: Enforce TLS for Inbound Email Delivery
Opportunistic SMTP TLS encrypts a great deal of server-to-server mail, but a sender can normally fall back when STARTTLS disappears or a certificate cannot be authenticated. MTA-STS gives a recipient domain a cached policy that supporting senders can use to resist downgrade and MX impersonation. It also creates a new availability dependency, so enforcement belongs at the end of a staged rollout.
How MTA-STS changes SMTP TLS delivery
RFC 8461 uses two publication points. DNS advertises an MTA-STS policy with a TXT record at _mta-sts.example.com, including v=STSv1 and an id that changes when policy content changes. HTTPS serves the text policy from https://mta-sts.example.com/.well-known/mta-sts.txt.
The policy contains version, mode, one or more mx patterns and max_age. In testing mode, compliant senders can report policy failures without enforcing delivery refusal. In enforce mode, a sender that has a valid cached policy should not deliver to an MX that fails the policy and TLS authentication; it queues and retries according to normal SMTP behavior.
Why enforcement can defer legitimate mail
The dangerous error is enabling enforcement before every advertised mail exchanger is ready. One forgotten disaster-recovery host with a wrong certificate, one policy pattern that misses a provider hostname or an HTTPS endpoint that serves the wrong content can defer legitimate mail. DNS, HTTPS, MX inventory, certificate lifecycle and mail operations must share an owner.
MTA-STS is also not end-to-end encryption and does not protect messages after receipt. It secures a supported SMTP hop against specific downgrade and impersonation risks. Senders that do not implement it will continue using their own SMTP TLS policy. DANE is another authenticated transport mechanism and has different DNSSEC requirements.
Deploy MTA-STS from inventory to enforcement
- Inventory every MX target. Resolve the domain repeatedly, include provider and disaster-recovery paths, and identify all certificate names, chains and renewal owners.
- Correct STARTTLS first. Make every intended MX advertise STARTTLS and present a publicly trusted, unexpired chain whose identity matches the policy MX pattern.
- Host the HTTPS policy. Serve only the policy at the required
mta-stshostname and well-known path over valid HTTPS. Keep it highly available and monitor content, not just status code. - Publish TLS-RPT. Create transport reporting before enforcement so certificate, policy and connection failures can be observed across participating senders.
- Start with mode testing. Publish exact
mxentries, a consideredmax_ageand testing mode. Increment the DNS policy ID whenever the served policy changes. - Exercise failure and recovery. Test each MX directly, certificate expiry alerting, host removal, provider failover, policy deployment, DNS updates and a documented rollback.
- Move to enforce deliberately. Require clean external validation and stable reporting across a meaningful traffic period. Obtain approval from mail, DNS and security owners.
- Operate the cached-policy lifecycle. Monitor HTTPS, DNS, MX inventory, certificate chain and TLS-RPT continuously. Remember that senders may retain a cached policy until max_age expires.
Choose testing, enforce or corrective action
| Condition | Safe mode | Operator decision |
|---|---|---|
| MX or certificate inventory is incomplete | Do not enforce | Finish discovery and assign lifecycle ownership |
| Policy is reachable and failures are still appearing | testing | Correct hosts and patterns; do not hide failures by ignoring reports |
| All MX paths validate and reporting is stable | candidate for enforce | Run failure drills and approve a change window |
| Enforce mode defers mail after a bad certificate deploy | enforce remains cached | Restore the valid certificate or serving path immediately |
| Emergency requires policy relaxation | Change policy and DNS id | Understand cached policies may delay relief; repair TLS is usually faster |
Worked rollout across two MX hosts
A domain uses mx1.example.net and mx2.example.net. Its policy lists both exact hosts and begins in testing mode. Reports reveal that the second server presents a certificate only valid for an internal load-balancer name. Ordinary opportunistic TLS had hidden the inconsistency because senders could continue without authenticated enforcement.
The team installs the correct chain on both paths, tests hostname validation externally and verifies automated renewal. It increments the policy ID, monitors another report period and then enables enforce mode. When a later deployment accidentally removes an intermediate certificate, alerts point to the affected host and the team restores the chain rather than weakening policy.
MTA-STS service and certificate evidence to retain
- DNS: one MTA-STS TXT record, expected version and ID, authoritative consistency, TTL and change history.
- Policy service: HTTPS status, certificate validity, exact content, MIME handling, latency, availability and checksum by region.
- MX fleet: resolved hosts and addresses, STARTTLS advertisement, certificate chain, SAN match, protocol and cipher policy.
- TLS-RPT: successful and failed sessions by reporter, result type, receiving host and policy string.
- Change evidence: policy revision, DNS ID, reviewer, deployment time, validation results, rollback steps and certificate owner.
MTA-STS deployment and recovery mistakes
- Enforcing before inventory: a low-priority or backup MX can still receive real delivery attempts.
- Serving a valid website with the wrong path: the policy must exist at the required well-known location and use the expected text format.
- Forgetting to change the DNS ID: supporting senders use it to discover that a cached policy should be refreshed.
- Using an MX pattern that does not match certificate identity: routing, policy and certificate names must be designed together.
- Expecting instant rollback: cached enforce policies remain relevant until their max_age, so restoring correct TLS is the primary recovery path.
MTA-STS enforcement checklist
- Enumerate primary, secondary, provider and disaster-recovery MX paths.
- Validate STARTTLS, public trust, expiration, complete chain and hostname on every host.
- Serve the policy from the exact HTTPS hostname and well-known path.
- Publish TLS-RPT and operate a safe aggregate-report pipeline.
- Begin with testing mode and precise MX entries.
- Increment the DNS policy ID whenever policy content changes.
- Run bad-certificate, unavailable-policy and MX-failover exercises before enforcement.
- Monitor continuously and keep certificate and policy rollback ownership current.
Publish both required MTA-STS components
DNS advertises that a policy exists and provides a change identifier. The HTTPS endpoint contains the policy itself.
_mta-sts.example.com. 3600 IN TXT "v=STSv1; id=2026082901"version: STSv1
mode: testing
mx: mx1.example.net
mx: mx2.example.net
max_age: 86400Serve the text file at https://mta-sts.example.com/.well-known/mta-sts.txt with a publicly trusted HTTPS certificate for mta-sts.example.com. The web hostname does not need to be an MX, and the policy is not served from the main website path.
Change the DNS id whenever policy content changes so supporting senders know to refresh. Use a deployment identifier that operators can map to version control and change records.
Match policy MX patterns to the real receiving fleet
| Policy entry | Matches | Does not match |
|---|---|---|
mx: mx1.example.net | That exact MX hostname | Another numbered host or an unrelated backup MX |
mx: *.example.net | Valid subdomains covered by the wildcard rule | The bare example.net name |
Every MX returned for the protected domain must match at least one policy pattern and present a certificate valid for its MX hostname under the RFC rules. Inventory provider failover, backup and disaster-recovery hosts before enforcement. A low-priority MX is still part of the delivery surface.
Validate the public MTA-STS path before enforcement
# Confirm one discovery record
dig +short TXT _mta-sts.example.com
# Retrieve the exact policy path
curl -i https://mta-sts.example.com/.well-known/mta-sts.txt
# Enumerate advertised MX hosts
dig +short MX example.com
# Inspect STARTTLS and the public certificate chain
openssl s_client -starttls smtp \
-connect mx1.example.net:25 -servername mx1.example.net \
-verify_return_error </dev/nullRepeat the TLS check for each resolved address behind every MX. Confirm STARTTLS is advertised after EHLO, the certificate is current, the chain is complete, and the identity matches. Test from outside the hosting network and from more than one resolver or region when split DNS or a global load balancer is involved.
Use a staged MTA-STS rollout with explicit gates
| Stage | Policy | Exit gate |
|---|---|---|
| Inventory | No published enforcement dependency | All MX, addresses, certificates and owners documented |
| Observe | TLS-RPT active | Report ingestion works and direct TLS checks agree |
| Testing | mode: testing with a bounded max_age | No unexplained multi-reporter policy failures |
| Failure drill | Testing remains active | Bad certificate, missing host and policy rollback procedures succeed |
| Enforce | mode: enforce | Change approval, monitoring and on-call ownership confirmed |
| Operate | Enforce with reviewed max_age | Continuous DNS, policy, MX and certificate checks |
Do not copy a long max_age from another organization during the first rollout. Choose it deliberately because senders may cache an enforce policy until it expires.
Recover from an enforced-policy incident without guessing
When valid senders defer mail under a cached enforce policy, restoring correct TLS is usually the fastest recovery. Identify the failing MX or certificate, redeploy the complete trusted chain, restore STARTTLS or correct the route. Removing the DNS discovery record or lowering max_age does not erase policies already cached by senders.
| Incident | Immediate repair | Follow-up |
|---|---|---|
| Expired certificate | Renew and deploy it to every listener | Test renewal automation and expiry alerts |
| Incomplete chain | Serve the required intermediate certificates | Validate each load-balancer backend |
| MX missing from policy | Restore intended routing or update policy and DNS ID | Repair inventory/change-control integration |
| STARTTLS absent on one address | Restore the listener or remove the broken address safely | Monitor EHLO capabilities per address |
| Policy endpoint unavailable | Restore highly available HTTPS service | Review caching behavior and multi-region monitoring |
Preserve TLS-RPT reports, SMTP deferral evidence, the policy version, DNS ID, deployment logs and repair time. After direct tests pass, monitor subsequent reports and inbound queue recovery.
Serve a minimal MTA-STS policy with Nginx
The policy hostname can use a small dedicated virtual host. Keep redirects, application routing and authentication away from the required well-known path.
server {
listen 443 ssl;
server_name mta-sts.example.com;
ssl_certificate /etc/letsencrypt/live/mta-sts.example.com/fullchain.pem;
ssl_certificate_key /etc/letsencrypt/live/mta-sts.example.com/privkey.pem;
location = /.well-known/mta-sts.txt {
default_type text/plain;
alias /var/www/mta-sts/mta-sts.txt;
}
location / { return 404; }
}Validate configuration before reload and retrieve the public file afterward.
sudo nginx -t
sudo systemctl reload nginx
curl -fsS https://mta-sts.example.com/.well-known/mta-sts.txtThe example assumes certificate automation and directory permissions are already managed. Apply the same availability monitoring used for the receiving mail service, and verify the response body rather than checking only HTTP 200.
Understand policy discovery, caching and delivery behavior
A supporting sender discovers the DNS record, retrieves a policy when needed and caches a valid policy for no longer than its declared max_age. During delivery it compares the selected MX with the cached policy and requires authenticated TLS when enforcement applies.
If a compliant sender has a valid enforce policy and cannot establish a policy-conforming TLS session, it should not fall back to cleartext delivery to that MX. The message remains queued and follows normal retry and expiry behavior. This protects against downgrade but means certificate and routing failures become mail-availability incidents.
Policy retrieval failure does not always have the same effect: behavior depends on whether the sender already has a valid cached policy. Operations must therefore assume different remote senders may hold different policy versions during a change. Preserve old and new policy content, DNS IDs and deployment times when diagnosing mixed reports.
Understand the two MTA-STS publication artifacts and their cache roles
MTA-STS publication has two coordinated parts. The TXT record at _mta-sts.example.com carries a changing policy identifier. The HTTPS endpoint at https://mta-sts.example.com/.well-known/mta-sts.txt carries the policy itself. A sender that notices a new identifier fetches and validates the HTTPS policy, then caches it for the policy max_age. Updating only one part creates behavior that varies by sender and cache age.
_mta-sts.example.com. 3600 IN TXT "v=STSv1; id=20260829-01"
version: STSv1
mode: enforce
mx: mx1.example.com
mx: *.mail-gateway.example.net
max_age: 604800The policy host is an HTTPS service, not an SMTP MX. Its publicly trusted certificate must cover mta-sts.example.com. Serve the exact well-known path without authentication, unrelated redirects or an HTML wrapper. Treat the identifier as a controlled release version, not a value that changes on every automation run.
Validate every MX, address and certificate against the policy
| Layer | Question | Failure consequence in enforce mode |
|---|---|---|
| MX DNS | Does the recipient domain resolve to an authorized policy pattern? | A cached-policy sender must not deliver to an unauthorized MX |
| SMTP capability | Does every target advertise STARTTLS after EHLO? | Delivery is deferred instead of falling back to clear text |
| Certificate identity | Does the SAN match the actual MX hostname? | Host validation fails even when encryption negotiates |
| Certificate chain | Is the chain current and publicly trusted? | Expired, incomplete or untrusted chains cause TLS failure |
| Load balancer | Do all nodes present the intended chain and capability? | Intermittent failures appear on particular addresses |
Test primary, equal-preference, backup and disaster-recovery MX records. An infrequently used backup with an expired certificate can become the production route during a primary failure and convert resilience into an outage. Record hostnames, addresses, preference, certificate fingerprint, SANs, issuer, expiry and STARTTLS result in one inventory.
Move from testing to enforce without creating a cached outage
- Inventory every advertised and standby MX plus systems that can change DNS, certificates and load-balancer membership.
- Deploy valid STARTTLS and certificates to all nodes before publishing a restrictive policy.
- Publish a testing policy with a modest cache period and enable TLS-RPT.
- Observe reports and independent probes across normal, failover and maintenance conditions.
- Correct host mismatches, incomplete chains, fetch failures and undocumented backup routes.
- Publish
mode: enforce, then change the TXT identifier when the HTTPS artifact is reachable. - Increase
max_ageonly after renewal, rollback and disaster recovery have been exercised.
mode: testing requests reporting from supporting senders without requiring enforcement. It does not test every route. Some senders do not report, reports are delayed, and normal traffic may not exercise a backup MX. Combine reports with active probes and receiving logs.
Design rollback around sender caches, not DNS intuition
MTA-STS is deliberately cacheable. If an enforced policy allows only mx1.example.com, changing DNS to an unlisted emergency MX can still fail for senders holding that policy. Lowering max_age during an incident does not shorten a cache stored under the former value. Removing the TXT record does not instantly remove a valid cached policy.
| Change | Safe preparation |
|---|---|
| Add a new MX | Add it to policy, release a new ID, allow cache propagation, then advertise it in MX DNS |
| Remove an MX | Drain it while service remains available during cache overlap |
| Rotate certificate | Deploy the complete trusted chain everywhere and probe before removing the former certificate |
| Emergency failover | Pre-authorize and continuously test disaster-recovery MX patterns |
| Policy-host outage | Use redundant HTTPS and DNS hosting |
The fastest safe recovery is commonly to restore compliant STARTTLS, a certificate or a previously authorized route. Weakening policy may take longer and expands downgrade exposure.
Model what a sender does with a valid cached policy
mx_hosts = dns_lookup_mx(recipient_domain)
policy = valid_cached_policy(recipient_domain)
?? fetch_policy_if_txt_id_changed(recipient_domain)
for mx in preference_order(mx_hosts):
if policy.mode == "enforce" and !policy.matches(mx.hostname): continue
session = smtp_connect(mx)
if policy.mode == "enforce":
require(session.starttls_advertised)
require(pkix_valid(session.certificate, mx.hostname))
attempt_delivery(session)
if no compliant route succeeds: queue_and_retryThis is explanatory pseudocode, not a complete RFC implementation. Temporary DNS, HTTPS and TLS errors require standards-compliant retry behavior. For receiving operators, the important point is that mail can remain queued while one manual probe succeeds because the sender may hold another cached policy or select a different MX address.
Troubleshoot an enforced-delivery incident without disabling security blindly
Start with the recipient domain, affected senders, first failure time and exact SMTP/TLS evidence. Query MX and MTA-STS TXT from authoritative and recursive resolvers. Fetch the well-known policy with certificate verification. Test every MX address using its hostname for validation, compare TLS-RPT failures by reporter and host, and inspect certificate or load-balancer changes at the same UTC time.
| Symptom | Likely boundary | Corrective action |
|---|---|---|
| Policy fetch errors from several reporters | HTTPS, DNS, CDN or certificate | Restore the exact public endpoint and verify independently |
| Host mismatch on one MX | Certificate SAN or virtual host | Deploy the intended certificate to every node |
| STARTTLS missing intermittently | Uneven MTA configuration | Remove or repair the noncompliant node |
| Mail queued after failover | Emergency MX absent from cached policy | Restore an authorized route and fix failover design |
Record remediation time and keep watching delayed reports. A report interval can include sessions from before the repair. Compare receiving logs with reporter windows, while recognizing that not every sender implements the same policy and reporting behavior. Close only after every advertised route is compliant, queued delivery recovers, certificate automation is corrected and fresh host-specific failures stop.
Assign MTA-STS ownership across DNS, HTTPS and mail operations
One service spans several teams. DNS owns the TXT release signal; web or CDN operations own the policy endpoint; messaging owns MX and STARTTLS; security owns certificate standards; incident response needs authority to restore every layer. Keep one change record with policy hash, DNS ID, certificate fingerprints, MX inventory and rollback owner. Monitoring must alert on unauthorized policy changes as well as outages.
Primary references
- RFC 8461: SMTP MTA Strict Transport Security
- RFC 8460: SMTP TLS Reporting
- RFC 7672: SMTP Security via DANE


