TLS-RPT for Email: Publish, Read and Act on Reports
TLS failures between mail servers are difficult to see from the receiving side. A certificate can expire, an MX host can stop advertising STARTTLS or a policy can reject a name mismatch while ordinary application monitoring remains green. SMTP TLS Reporting, or TLS-RPT, lets domains request aggregate reports from participating sending systems.
What TLS-RPT observes in SMTP transport
RFC 8460 defines a TXT policy at _smtp._tls.<policy-domain>. Version 1 begins with v=TLSRPTv1 and uses rua to name one or more report destinations. Supported URI forms include mailto: and https:. Reports use an I-JSON schema and describe policy, totals and failure details for a reporting period.
TLS-RPT is observability, not enforcement. Publishing it does not require a sender to use TLS and does not make a receiver’s certificate valid. MTA-STS or DANE can express transport policy; TLS-RPT reveals how participating senders experienced that policy or the receiving TLS service.
Why TLS reports need secure interpretation
A report endpoint can itself become an operational and security risk. Mailbox delivery produces compressed MIME attachments; HTTPS accepts remote submissions. Parsers must limit body and decompressed size, reject unsafe archive paths, validate JSON structure, retain raw evidence securely and avoid executing report content.
Individual failure counts also need context. A transient outage can affect one reporting organization, while a bad certificate deployed to one MX host creates a distributed pattern. Compare failure type, MX host, policy string, reporting organization and time window. Do not treat the absence of reports as proof that all mail uses TLS because participation is not universal.
Deploy a TLS-RPT reporting pipeline
Roll out the reporting path as an owned service, with documented security and retention boundaries.
- Inventory the receiving path. List every advertised MX host, certificate name, issuing chain, STARTTLS behavior, gateway provider and operational owner.
- Create a dedicated endpoint. Use a purpose-specific mailbox or HTTPS service with authentication and abuse controls appropriate to the chosen URI method.
- Publish one valid TXT policy. Create
_smtp._tls.example.com TXT "v=TLSRPTv1; rua=mailto:[email protected]". Avoid multiple competing TLS-RPT TXT records. - Authorize external destinations. When the report destination is outside the policy domain, implement the external reporting authorization required by RFC 8460.
- Ingest defensively. Enforce transport, content-type, compression, archive, JSON-depth, record-count and storage limits before accepting data into analytics.
- Normalize without losing evidence. Retain report ID, organization, date range, policy, MX host, result type and counts while preserving the raw report checksum.
- Alert on actionable patterns. Use sustained or distributed failures by host and result type. Suppress duplicates and route certificate, DNS, network and provider faults to the correct owner.
- Verify the fix externally. Test DNS, STARTTLS, certificate chain and hostname from more than one network, then watch later report windows for recovery.
Map TLS-RPT failure patterns to investigations
| Report pattern | Likely investigation | First action |
|---|---|---|
| certificate-expired across many reporters | Certificate lifecycle or stale MX node | Inspect every advertised MX and deploy the complete renewed chain |
| starttls-not-supported on one host | Listener or load-balancer inconsistency | Compare SMTP capability responses across all addresses |
| certificate-host-mismatch | Certificate names do not cover the policy MX pattern | Validate policy mx entries and SAN names together |
| dnssec-invalid for a DANE policy | DNSSEC or TLSA chain problem | Use authoritative DNS diagnostics before changing mail policy |
| Small isolated connection failures | Network path or reporter-specific condition | Correlate timestamps and avoid declaring a global incident |
Worked incident: one MX serves an expired certificate
A receiving domain has two MX hosts behind separate load balancers. A renewed certificate is deployed to the primary path, but the secondary continues serving the expired chain. Uptime checks against the primary pass. TLS-RPT shows certificate-expired failures tied to the secondary MX from several independent reporting organizations.
Operators confirm the old listener, deploy the correct certificate and test both hostnames directly with STARTTLS. They do not delete or weaken the MTA-STS policy to hide the symptom. Later report windows show successful sessions and the failure pattern disappears, providing external confirmation that the whole receiving fleet is consistent.
TLS-RPT fields and operational evidence to retain
- Policy identity: policy domain, policy type, policy strings, policy version or ID and the MX pattern evaluated.
- Report identity: reporting organization, report ID, contact information, date range and ingestion time.
- Session totals: successful-session count and failed-session count, trended by policy and reporter.
- Failure detail: result type, receiving MX hostname, receiving IP, failed-session count and additional information.
- Pipeline health: endpoint errors, rejected payloads, parse failures, duplicates, lag, raw checksum and storage retention.
TLS-RPT publishing, parsing and alerting mistakes
- Assuming reports enforce TLS: they describe observations; transport policy is a separate mechanism.
- Publishing more than one policy record: ambiguity can cause the domain to be treated as not implementing TLS-RPT.
- Ignoring external-destination authorization: reporters may decline to send reports to an unauthorized third-party domain.
- Extracting compressed reports unsafely: validate paths and size to prevent overwrite and decompression attacks.
- Alerting on every count: aggregate reports arrive after the fact and need correlation, deduplication and materiality rules.
TLS-RPT deployment checklist
- Inventory every MX host and verify STARTTLS, hostname and certificate chain.
- Choose a controlled mailto or HTTPS aggregate-report endpoint.
- Publish exactly one syntactically valid TLSRPTv1 TXT policy.
- Authorize an external destination when the reporting domain differs.
- Limit payload, decompressed data, JSON complexity and retention.
- Normalize report identity, policy, totals, failure type and receiving host.
- Alert on sustained multi-reporter patterns and route by fault domain.
- Verify remediation directly and in subsequent aggregate report windows.
Publish the TLS-RPT DNS record and external authorization
For reports delivered to the same organizational domain, publish one TXT record at the TLS-RPT policy name.
_smtp._tls.example.com. 3600 IN TXT \
"v=TLSRPTv1; rua=mailto:[email protected]"If example.com sends reports to reports.vendor.example, the destination domain must authorize that use. RFC 8460 defines the authorization lookup under the destination domain.
example.com._report._smtp._tls.reports.vendor.example. 3600 IN TXT \
"v=TLSRPTv1"Publish through the authoritative DNS service and verify the response from more than one resolver. Multiple TLS-RPT policy records can make the policy invalid; do not create separate TXT records for separate destinations.
Read the structure of an aggregate TLS report
A report covers a time interval and groups successful and failed sessions by policy. The shortened example below demonstrates the main fields; production reports may contain several policy and failure-detail objects.
{
"organization-name": "Sending Operator",
"date-range": {
"start-datetime": "2026-08-28T00:00:00Z",
"end-datetime": "2026-08-28T23:59:59Z"
},
"contact-info": "[email protected]",
"report-id": "2026-08-28-example.com",
"policies": [{
"policy": {
"policy-type": "sts",
"policy-domain": "example.com",
"mx-host": ["mx1.example.com"]
},
"summary": {"total-successful-session-count": 940,
"total-failure-session-count": 12},
"failure-details": [{
"result-type": "certificate-expired",
"receiving-mx-hostname": "mx1.example.com",
"failed-session-count": 12
}]
}]
}Use the report ID, reporting organization, date range and raw payload checksum for deduplication. Counts are aggregate observations from that reporter, not individual-message logs.
Route TLS-RPT failure types to the correct owner
| Result family | Likely fault domain | Verification |
|---|---|---|
| starttls-not-supported | SMTP listener, proxy or inconsistent backend | Compare EHLO capabilities on every advertised address |
| certificate-expired | Certificate deployment or forgotten MX | Inspect validity dates and chain on every listener |
| certificate-host-mismatch | Certificate SAN or MTA-STS MX pattern | Compare policy match and authenticated hostname rules |
| certificate-not-trusted | Incomplete chain or trust path | Retrieve the served chain externally, including intermediates |
| validation-failure | Policy retrieval or policy validation | Check policy content, DNS ID and HTTPS endpoint |
| dnssec-invalid | DNSSEC/TLSA path for DANE | Validate the authoritative chain and TLSA records |
Preserve the exact result-type and any additional information. Do not collapse every failure into “TLS failed”; the remediation owners are different.
Verify DNS, STARTTLS and certificates directly
# TLS-RPT policy
dig +short TXT _smtp._tls.example.com
# Advertised MX fleet
dig +short MX example.com
# STARTTLS and the certificate served by one MX
openssl s_client -starttls smtp \
-connect mx1.example.com:25 -servername mx1.example.com \
-showcerts </dev/null
# MTA-STS policy, when used
curl -fsS https://mta-sts.example.com/.well-known/mta-sts.txtRun checks from an external system because an internal path may bypass the public listener, certificate or DNS view. Test every MX hostname and resolved address. Report recovery is delayed by reporting intervals, so direct validation proves the immediate repair while later aggregate windows confirm distributed behavior.
Treat the TLS-RPT endpoint as an untrusted ingestion service
| Boundary | Required control |
|---|---|
| Mailbox or HTTPS request | Rate limits, accepted content types, maximum body size and authentication where the transport supports it |
| Compression | Compressed and expanded size limits; reject nested or unexpected archives |
| JSON parsing | Schema, depth, string length, array-count and numeric-range validation |
| Storage | Raw checksum, restricted access, retention limit and encrypted backups |
| Alerting | Deduplicate by reporter and report ID; require material duration or independent reporters |
Do not render report-controlled strings as trusted HTML or use them in shell commands. Keep parsing failures visible: a sudden rise may indicate an upstream format change, abuse, or a regression in the ingestion service.
Choose mailto or HTTPS reporting with operational ownership
A mailto: destination is easy to publish and can use an existing secure mail path, but the consumer must correctly parse MIME messages and supported compressed attachments. Isolate the mailbox from ordinary correspondence, prevent autoresponders and monitor quota, rejection and processing lag.
An https: destination avoids mailbox MIME handling but exposes a public ingestion endpoint. Require HTTPS, strict request-size and processing limits, safe JSON handling, rate control and a response strategy that does not leak internal details. Monitor certificate renewal and regional reachability because an unavailable report service creates an observability gap.
| Question | Mailto endpoint | HTTPS endpoint |
|---|---|---|
| Primary parser | MIME plus compression plus JSON | HTTP request plus compression/content encoding plus JSON |
| Capacity risk | Mailbox quota and attachment controls | Request flood and worker exhaustion |
| Abuse control | Mail filtering without discarding valid reports | Rate, body-size and concurrency limits |
| Availability evidence | Accepted mail and processing lag | External HTTPS probes and accepted-report counts |
Alert on material patterns instead of raw failure counts
Build a baseline by reporter, protected domain, policy type and receiving MX. A count of ten failures can be severe for a host that normally receives ten sessions and insignificant in a fleet with millions of successful sessions. Review both absolute failures and failure ratio, while protecting against tiny denominators.
- Page immediately when several independent reporters show certificate expiry or hostname mismatch on the same active MX.
- Create a ticket for sustained low-volume failures that repeat across reporting periods.
- Route DNSSEC and DANE failures to DNS ownership rather than certificate operations.
- Alert separately when reports stop arriving from established reporters or the parser rejection rate rises.
- Suppress exact duplicate report IDs but retain conflicting payloads as a security or integration anomaly.
TLS-RPT is delayed aggregate telemetry. Do not use it as the only real-time check. Pair it with active MX, STARTTLS, policy and certificate monitoring.
Publish and verify the TLS-RPT policy record
_smtp._tls.example.com. 3600 IN TXT
"v=TLSRPTv1; rua=mailto:[email protected]"The policy belongs at _smtp._tls under the receiving domain. Use a dedicated address or HTTPS report collector supported by the specification and service design. When the report destination is outside the policy domain, implement the required external-destination authorization rather than assuming third-party delivery will work.
Query every authoritative server and public resolvers, check TXT concatenation, and retain the exact record and TTL. A syntactically valid record does not prove the reporting mailbox, decompression or parser path works.
Normalize the report without discarding raw evidence
{
"organization-name": "Example Reporter",
"date-range": {"start-datetime": "...", "end-datetime": "..."},
"contact-info": "[email protected]",
"report-id": "2026-08-29-example",
"policies": [{
"policy": {
"policy-type": "sts",
"policy-string": ["version: STSv1", "mode: enforce", "..."],
"policy-domain": "example.com",
"mx-host": ["mx1.example.com"]
},
"summary": {
"total-successful-session-count": 182400,
"total-failure-session-count": 27
},
"failure-details": []
}]
}Store the original MIME attachment and compressed payload checksum, reporter, report ID, date range, policy identity, summary counts, failure rows and parser version. Report IDs are reporter-scoped; deduplicate with reporter and date range, not globally by ID alone.
Map RFC 8460 result types to the responsible layer
| Result type | Likely boundary | First check |
|---|---|---|
starttls-not-supported | MX SMTP capability | EHLO on the reported MX and route |
certificate-host-mismatch | Certificate identity | SAN against policy MX pattern |
certificate-expired | Certificate lifecycle | Active node, SNI and deployment time |
certificate-not-trusted | PKIX chain/trust | Chain, CA, intermediates and reporter reason |
sts-policy-fetch-error | HTTPS/DNS policy publication | Well-known URL, DNS, CDN and TLS |
sts-policy-invalid | Policy syntax/semantics | Served text and policy ID release |
tlsa-invalid or dnssec-invalid | DANE/DNSSEC | Validated TLSA chain and zone state |
validation-failure | Unclassified validation | Failure reason, reporter and direct probe |
Interpret aggregate counts and overlap cautiously
RFC 8460 notes that failure types are non-exclusive. One failed session can contribute to more than one failure detail, so summing every detail can exceed the total failed-session count. Reporters also differ in volume, policy support and timing. Do not calculate a global failure percentage by blindly summing heterogeneous reports.
report_failure_rate =
total_failure_session_count
/ (total_successful_session_count + total_failure_session_count)
do_not_assume:
sum(failure_detail.failed_session_count)
== total_failure_session_countSegment by reporter, policy domain, MX host, result type and date. Compare persistent failures across independent reporters with direct probes and receiving MTA logs.
Treat reports as untrusted external input
| Risk | Control |
|---|---|
| Compression bomb | Compressed and expanded size, ratio and recursion limits |
| Malformed JSON | Strict parser, schema validation and quarantine |
| Duplicate flood | Reporter/report-ID/date idempotency and rate limits |
| Forged reporter identity | Preserve transport evidence; do not treat Contact-Info as authentication |
| Unsafe URI or text | Escape presentation; never auto-fetch arbitrary failure URLs |
| Personal/internal data exposure | Access controls, minimization and retention policy |
Parse in an isolated worker without shelling out to user-controlled filenames. Store unknown result types for forward compatibility and alert on parser rejection rates.
Correlate reports with policy, certificates and inbound service
A certificate deployment to one backup MX can create intermittent host-mismatch reports while normal probes against the primary remain green. Group by reported MX host, then test every address behind that host with the correct SNI and policy pattern. Compare certificate deployment times and load-balancer membership.
An sts-policy-fetch-error spike can result from DNS, CDN, HTTPS certificate or path behavior even while SMTP is healthy. Preserve the served policy and response headers from multiple networks. During enforced MTA-STS failure, restore compliant TLS and policy availability first; lowering policy mode or max-age does not instantly clear sender caches.
Build a TLS-RPT operations dashboard with explicit unknowns
| Panel | Dimensions | Alert condition |
|---|---|---|
| Report intake | Reporter, report date, arrival delay | Expected reporters disappear or parser rejects rise |
| Session summary | Policy domain and reporter | Persistent failure-rate deviation with minimum sessions |
| Failure taxonomy | Result type and MX host | New or concentrated certificate/policy failure |
| Certificate fleet | MX, address, SAN and expiry | Mismatch, expiry window or inconsistent chain |
| MTA-STS publication | DNS ID, HTTPS policy hash and mode | Unexpected change or fetch failure |
Missing reports mean unknown coverage, not zero failures. Keep direct external STARTTLS tests and receiving service monitoring alongside aggregate reporter evidence.
Document TLS-RPT blind spots before interpreting silence
| Blind spot | Operational consequence |
|---|---|
| Non-participating senders | Their TLS successes and failures never appear |
| Transient network failures not reported | Internal service monitoring may show incidents absent from aggregates |
| Delayed or missing aggregate delivery | Current-day dashboard is incomplete |
| Reporter-specific policy support | Different senders exercise MTA-STS, DANE or opportunistic TLS differently |
| Non-exclusive failure types | Detail counts cannot be treated as disjoint sessions |
| Aggregated source information | One report cannot identify every affected sender or message |
Maintain a reporter coverage inventory and a data-freshness state. Green with no recent reports is unknown, not healthy. TLS-RPT complements certificate expiry monitoring, DNS/HTTPS probes, SMTP capability tests and receiving MTA logs.
Establish expected arrival ranges for reporters that normally contribute material traffic, but do not treat a missing report as proof that the reporter stopped sending. Check the reporting mailbox, HTTP collector, authentication, decompression, parser quarantine and storage pipeline first. Compare aggregate session volume with receiving MTA connections only as a directional coverage check because the populations and policy support differ.
Run a controlled certificate and policy release test
- Enumerate every primary, equal-preference, backup and disaster-recovery MX host and address.
- Probe EHLO, STARTTLS, certificate chain, SAN and protocol from independent networks.
- Fetch the MTA-STS policy at the exact well-known path and compare its hash with the reviewed artifact.
- Query the DNS policy ID and TLS-RPT record from authoritative and recursive resolvers.
- Introduce a controlled lab mismatch or unavailable host outside production and confirm the parser classifies the expected result.
- Deploy one bounded production node, monitor direct probes and inbound sessions, then expand.
- Verify subsequent reports by reporter and MX host while respecting report delay.
Keep the former certificate and configuration available for precise rollback. Do not solve one bad node by deleting it from DNS before understanding cached MX and policy behavior.
Record the policy DNS ID, policy file checksum, active certificate fingerprints, load-balancer membership and exact deployment time in UTC. A later report covers a time interval and can include sessions before and after the release. Without this timeline, the team can incorrectly conclude that a fixed certificate is still failing or that a delayed report describes the current fleet.
Primary references
- RFC 8460: SMTP TLS Reporting
- RFC 8461: SMTP MTA Strict Transport Security
- RFC 7672: SMTP Security via DANE


