Email Header Analysis: Trace Routing, Authentication and Delays
A raw email header is a sequence of claims added by different systems. Some fields are supplied by the sender, while receiving systems prepend transport evidence. Analysis must begin at a trusted boundary, read the Received chain from the bottom upward, and distinguish a cryptographic result from a displayed address. One suspicious line rarely proves spoofing; the conclusion comes from routing, identity, authentication, timing and message context together.
Capture the complete original safely
Use the mailbox client’s “show original” or equivalent export and retain the full MIME message when authorization permits. Forwarding or copying visible fields can remove Received hops, Authentication-Results, ARC, boundaries and encodings. Calculate a checksum, record acquisition time, mailbox, folder and collection method, and work on a copy.
Headers contain addresses, IPs, internal hostnames and tracking identifiers. Restrict access, redact only in presentation copies, and preserve the unmodified evidence under the applicable retention policy.
Establish the trusted header boundary
Start with the topmost Received and Authentication-Results fields added by infrastructure you trust. Fields below that boundary may describe earlier hops but can be forged before the message enters a trusted system. Read Received fields from bottom to top to reconstruct the apparent path, then compare each receiving host, source address, protocol, TLS notation and timestamp.
Private IPs can be normal inside an organization. A missing public hop can result from internal architecture. Treat discrepancies as investigation leads, not automatic verdicts.
Compare every email identity
| Field | Role | Question |
|---|---|---|
| From | Visible author identity | Does the user-facing domain match expectation? |
| Return-Path | Final envelope sender recorded at delivery | Does SPF identity align for DMARC? |
| DKIM d= | Signing domain | Did a trusted verifier pass it and does it align? |
| Reply-To | Reply destination | Is a different destination legitimate? |
| Message-ID | Message identifier | Does domain/format fit the sending system? |
A different return path is common with ESPs; DMARC alignment and authorized configuration determine whether it is acceptable.
Read Authentication-Results in context
Trust results written by the receiving system in its documented authentication-results namespace. SPF evaluates the connecting IP against the envelope identity. DKIM verifies signed headers/body with the selector key. DMARC passes when a passing SPF or DKIM identifier aligns with the visible From domain. ARC can preserve an intermediary’s authentication assessment, but an ARC pass is not proof that the message is safe.
Authentication-Results: mx.example;
spf=pass smtp.mailfrom=bounces.example.org;
dkim=pass header.d=example.org header.s=s1;
dmarc=pass header.from=example.orgBuild a normalized timeline
Convert timestamps to UTC while retaining the original zone and text. Compare adjacent Received hops, Date, delivery time and any queue identifiers. Clock skew, time-zone formatting and delayed queues can explain small reversals. Large impossible jumps, a future Date or an unrecognized relay require corroboration.
Map hostnames to DNS and ownership only as they existed or were observed at investigation time. Avoid geolocating privacy relays or cloud addresses as if they prove the sender’s physical location.
Separate anomalies from verdicts
- Visible From and aligned authenticated domain disagree with the expected organization.
- Reply-To points to an unrelated domain in a credential or payment request.
- Received path enters through an unexpected network or bypasses the secure gateway.
- DKIM fails after body modification, or no expected signature is present.
- Unicode lookalikes, misleading display names or mismatched link domains appear.
Combine header findings with message body, URLs, attachments, account telemetry and user report. Never click an untrusted URL merely to investigate it.
Unfold and parse headers without destroying evidence
RFC message fields can continue on following lines that begin with whitespace. Parse with a standards-aware MIME library and keep the raw octets. Do not split naively on every colon or comma: date values, group addresses, parameters and authentication properties contain delimiters. Duplicate fields can be legal, suspicious or security-relevant depending on field semantics and signer coverage.
raw_message_sha256: 8c3e...f21a
acquired_at_utc: 2026-08-29T11:42:18Z
source: mailbox export / Show Original
analysis_copy: immutable raw + separately parsed representation
preserve:
raw field order, folding, line endings, MIME boundaries,
complete body, attachment hashes and collection metadataWork from the complete original EML where authorization permits. A forwarded copy inserts a new envelope and often removes or transforms the evidence. Redact only an analyst presentation copy; the evidence checksum must correspond to the unmodified artifact.
Reconstruct Received hops from the trusted boundary downward
Each compliant relay prepends a Received field, so the top field is the most recent hop and the bottom is the earliest recorded claim. Begin with infrastructure controlled by the receiving organization. Read downward only as far as the trust model supports. An attacker can inject fake earlier Received lines before entering the first trusted gateway.
| Element | Question |
|---|---|
from | What identity/address did this receiving MTA report? |
by | Which system added the field? |
with | Which transport/protocol and TLS notation? |
id | Can the queue identifier be correlated? |
for | Which envelope recipient was recorded? |
| date | How does the timestamp compare with adjacent trusted hops? |
Private IPs and internal hostnames are normal inside trusted organizations. Public geolocation is weak evidence, especially for cloud and privacy relay networks.
Map SPF, DKIM, DMARC and ARC to the identities they actually evaluate
| Mechanism | Input identity | What pass establishes |
|---|---|---|
| SPF | Connecting IP plus MAIL FROM/HELO domain | IP authorized for that identity at evaluation time |
| DKIM | d=, selector and signed message portions | Signature verified with domain key |
| DMARC | Visible From aligned with passing SPF or DKIM | At least one aligned authentication path passed |
| ARC | Signed intermediary authentication chain | Chain integrity under validator result, not message safety |
Use Authentication-Results written by the trusted receiver’s documented authserv-id. A message can contain attacker-supplied fields with convincing names. A vendor DKIM pass may be unaligned with the visible From and therefore not contribute to DMARC.
Read the DKIM signature and failure reason precisely
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed;
d=example.com; s=mail2026;
h=from:to:subject:date:message-id;
bh=BASE64_BODY_HASH; b=BASE64_SIGNATUREd= identifies the signing domain, s= selects the DNS key, h= lists signed header fields and c= defines canonicalization. Verification failure can result from an unavailable key, revoked selector, body change, signed-header mutation, malformed syntax or unsupported policy. Preserve the verifier’s exact reason.
Check duplicate signed fields carefully. DKIM selects header instances according to the signature rules; adding another From or Subject can produce confusing displays and security issues even when a signature technically validates. Do not re-run verification on a message altered by ticketing or forwarding and treat the result as original.
Normalize time while retaining original values and uncertainty
Convert each trusted timestamp to UTC but store original text and offset. Calculate hop-to-hop delay, queue delay and the gap between the Date field and first trusted receipt. Small reversals can arise from clock skew. A forged Date field has little evidentiary weight compared with trusted server timestamps.
hop_timeline(
message_hash, hop_index, by_host, from_claim,
observed_source_ip, protocol, tls_detail,
original_timestamp, normalized_utc,
trust_class, queue_id
)Correlate queue IDs and provider logs where available. DNS and IP ownership can change, so record lookup time and source. Avoid asserting that an IP proves a person’s location or device. The timeline supports routing and delay conclusions, not author identity by itself.
Separate visible author, sender, reply and envelope identities
The displayed name is free text and can imitate a trusted brand. The RFC 5322 From address is the DMARC identity. Sender can identify an agent in some cases; Reply-To controls where a client sends replies; Return-Path records the final envelope sender after delivery. Message-ID is a correlation clue, not authentication.
| Pattern | Interpretation |
|---|---|
| Brand display name, unrelated From domain | Potential impersonation; check expected domains and context |
| Aligned From/DKIM, different ESP return path | Can be normal when aligned DKIM passes |
| Reply-To changed to unrelated domain | Investigate business process and request context |
| Message-ID domain differs | May reflect sending software; weak alone |
| Several From fields | Parsing/security anomaly requiring raw analysis |
Use Unicode-aware display and domain analysis for lookalikes, but do not decide solely from visual similarity.
Use ARC to explain forwarding without treating it as an allowlist
Forwarding can break SPF because the forwarder connects from another IP. Mailing lists can modify subject or body and break original DKIM. ARC lets an intermediary sign what it observed and bind that assessment into a chain. Validate set completeness, contiguous instances, ARC-Message-Signature, seals and computed chain result.
A cryptographically valid chain proves signed assertions were not altered; it does not prove every sealer is trustworthy. Compare sealer domain, known forwarding path, original authentication assertion, final DMARC and receiver policy. Unknown but valid sealers do not automatically rescue a DMARC failure.
When investigating, retain instance count, failing instance/field, selectors, key results and exact mutations. Test the original, intermediary output and final received artifact when available.
Combine header findings with content and account evidence safely
Header analysis can show an unexpected route, authentication failure, unrelated Reply-To or mismatched identity. It cannot prove a request is legitimate merely because DMARC passes; compromised accounts and attacker-owned domains authenticate. Examine URLs, attachment hashes, language, transaction context, account sign-ins and known vendor workflows.
| Finding | Next action |
|---|---|
| Authentication fails for claimed brand | Quarantine/escalate and verify through trusted channel |
| Authentication passes attacker-owned lookalike | Analyze domain age/context, content and account telemetry |
| Expected account sent unusual request | Investigate compromise and revoke sessions |
| Link domain differs from From | Resolve redirect chain in an isolated authorized system |
Do not click links or open attachments on an analyst workstation. Use sanctioned sandboxing and preserve chain of custody.
Worked header case: an aligned message is still malicious
A payment-change message displays a supplier’s real name and passes SPF, DKIM and DMARC for supplier-payments.example. Users assume pass means trustworthy. The organization normally uses supplier.example; the passing domain is a lookalike registered by the attacker.
The trusted Received chain shows a commodity hosting provider, the Reply-To uses another new domain, and account/procurement records contain no request. Authentication correctly establishes the attacker controls the lookalike domain and message; it does not establish supplier authorization.
Security blocks the lookalike domains, searches for related messages and confirms payment details through the known supplier channel. The report states exact authenticated identities and routing facts instead of saying “spoofed email” generically. This distinction guides domain monitoring and user training.
Write a defensible analyst report
- Evidence hash, source, acquisition method and UTC time.
- Trusted boundary and systems considered authoritative.
- Normalized hop timeline with raw references.
- SPF, DKIM, DMARC and ARC identities/results.
- Visible From, Reply-To, envelope and link-domain comparison.
- Content/account corroboration and stated limitations.
- Conclusion separated into observed fact, inference and unknown.
- Containment or follow-up with named owner.
Never paste full headers containing personal or internal data into public tools without authorization. Use local parsers or approved services and redact presentation copies only after evidence preservation.
Analyze MIME structure when authentication or display differs
The header block and body form one MIME artifact. Content-Type, transfer encoding, multipart boundaries and Content-Disposition determine what users and scanners see. A DKIM body-hash failure can arise when a gateway re-encodes MIME, adds a footer or normalizes line endings. Preserve the raw body and compare changes across hops.
Content-Type: multipart/alternative; boundary="b1"
--b1
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: quoted-printable
...
--b1
Content-Type: text/html; charset=UTF-8
Content-Transfer-Encoding: quoted-printable
...
--b1--Check whether HTML and text parts express the same intent, whether attachments have misleading names/types, and whether an outer message/rfc822 wrapper contains another header set. Do not treat headers from an attached message as transport evidence for the outer delivery.
Interpret mailing-list and automated-message fields in context
| Field | Operational use | Caution |
|---|---|---|
| List-ID | Stable list identity | Claimed field; validate expected system |
| List-Unsubscribe/Post | Standards-based unsubscribe metadata | Does not authenticate sender |
| Auto-Submitted | Signals automatic response | Useful for loop prevention, not security verdict |
| Precedence | Historical bulk/list hint | Nonstandard policy varies |
| Feedback-ID | Provider/campaign attribution | Format is ecosystem-specific |
Unexpected absence or value can help diagnose a sending integration, but attacker-controlled messages can add these fields. Anchor conclusions in the trusted path and authenticated identities.
Treat X-headers and vendor fields as implementation-specific
X- fields, spam scores, gateway verdicts and provider trace IDs can be useful inside the system that defines them. Document the generating product, trusted boundary, version and value semantics. Do not apply one vendor’s score thresholds to another or publish internal rule names broadly.
Trace identifiers can correlate logs, but they may contain tenant, campaign or recipient information. Protect them and never paste them into unapproved public analyzers. If a field appears below the first trusted hop, treat it as a sender claim until corroborated.
When a security gateway rewrites Subject or adds warning banners, determine whether it signs/seals after modification and whether the change explains downstream DKIM failure.
Validate parser output against raw fixtures
Create fixtures for folded fields, duplicate From/Subject, comments, quoted strings, encoded words, IPv6 Received clauses, malformed dates, several Authentication-Results and nested messages. Assert that the parser preserves order and exposes errors instead of silently repairing hostile syntax.
analysis_result(
raw_hash, parser_name, parser_version,
trusted_boundary_index, parse_warnings,
normalized_identities, auth_results,
timeline, analyst_conclusion_version
)Compare at least one independent standards-aware parser for security-sensitive cases. A convenient online analyzer can render diagrams but should not become the evidence authority, especially when confidentiality forbids uploading the message.
Use a repeatable header incident sequence
- Acquire and hash the full original under authorization.
- Identify the top trusted receiver fields and authserv-id.
- Reconstruct the route and UTC timeline.
- Map visible, envelope, DKIM, DMARC, ARC and reply identities.
- Inspect MIME, URLs and attachments in an isolated workflow.
- Correlate queue, account and security telemetry.
- State observed facts, supported inference and unknowns separately.
- Contain domains, accounts or credentials based on corroborated evidence.
Retain parser versions and lookup times. DNS, WHOIS-style registration data and cloud ownership observed today may differ from delivery time.
Resolve disagreement between header fields and a local verifier
A received header can say DKIM passed while a later local check fails because DNS keys changed, the saved message was altered, or the receiver evaluated before a selector was removed. Record who wrote Authentication-Results, when the message arrived and whether the artifact is byte-identical. A current DNS lookup cannot always reproduce historical verification.
| Disagreement | Investigation |
|---|---|
| Trusted receiver pass; local key missing | Selector retirement and receipt time |
| Trusted receiver pass; local body hash fails | Evidence transformation after delivery |
| Sender-supplied pass; trusted receiver fail | Ignore untrusted claim; use receiver result |
| Two trusted gateways differ | Mutation between gateways and auth timing |
State which result supports the conclusion and why. Never overwrite the original header with a freshly generated result.
Correlate headers with receiving infrastructure logs
correlation keys:
receiving queue ID
Message-ID plus recipient and time window
source IP / TLS session ID
gateway trace ID
final delivery ID
avoid relying on Message-ID alone: duplicates and attacker values existLogs can confirm the observed source address, SMTP envelope, recipient decision, TLS properties, content-filter result and downstream delivery. Use several keys and a tight UTC window. Protect logs because they reveal internal topology and recipients.
If logs conflict with the exported message, validate mailbox/client transformations and collection method. Document missing retention rather than inventing a hop.
State what header analysis cannot prove
Headers cannot prove a person authored a message, that authenticated content is benign, that a displayed brand authorized an attacker-controlled lookalike, or that an IP is a physical location. They may not expose internal provider decisions or every pre-submission system. Encryption-in-transit notation does not mean end-to-end confidentiality.
A defensible conclusion combines headers with account authentication, sending application logs, domain ownership at the relevant time, transaction context and isolated content analysis. Use probability language for inference and preserve unknowns.
Worked case: a twelve-hour delay inside the receiving path
A customer reports that an authenticated order message arrived twelve hours late. The sender MTA log shows immediate Gmail acceptance, so the campaign dashboard calls it delivered. The trusted Received chain shows the message entered the first receiving gateway at send time, remained queued at an internal policy gateway, and reached the mailbox hours later.
Queue identifiers correlate with a content-scanner outage. SPF, DKIM and DMARC passed, and no outbound retry caused the delay. The incident report assigns the failure to post-acceptance receiving processing, not sender transport or inbox placement. This conclusion is possible only because timestamps were normalized and trusted hops were separated from sender-supplied fields.
Final analyst check
Before closing, verify that every stated identity comes from the correct field, every authentication result comes from a trusted evaluator, and every time comparison retains its original offset. Link conclusions to raw evidence locations and record unresolved gaps explicitly.


