Format and Structure of an E-mail Message
Open “Show original” or export an EML file and the message looks nothing like the mail-client view. The raw source contains the delivery trace, visible identities, authentication results, content rules, alternative bodies, inline resources, and attachments. Once you know where one section ends and the next begins, many delivery and rendering problems stop looking mysterious.
The SMTP envelope is not the message
Start by separating transport data from message data. During SMTP, the sending system provides an envelope sender with MAIL FROM and one or more envelope recipients with RCPT TO. Mail servers use those values to route the message and send delivery status notifications.
The visible From, To, and Cc fields are part of the message header section. They can differ from the SMTP envelope. For example, a delivery platform may use a dedicated bounce domain as the envelope sender while showing the organization domain in the visible From field. Bcc recipients are normally represented in the SMTP envelope but omitted from the delivered message header section.
This distinction matters during authentication and bounce analysis. SPF evaluates the SMTP identity. DMARC then checks whether an authenticated SPF or DKIM domain aligns with the visible From domain. A mail client screen does not expose enough information to evaluate that path.
The top-level message structure
RFC 5322 defines an Internet message as a header section followed optionally by a body. One empty line separates the two sections.
Header-Field: value
Another-Field: value
Message body begins here.
Header fields use a field name, a colon, and a field body. Field names are case-insensitive, although familiar capitalization improves readability. A long field can be folded across physical lines by placing whitespace at the start of each continuation line. Parsers unfold it before interpreting the value, but an archive or forensic workflow should preserve the original raw representation.
Header fields are not guaranteed to appear in one fixed order. Some can occur more than once. Trace fields such as Received need special care because their order describes the delivery path.
Common Headers (A = Always Present, F = Frequent, O = Optional)
The A, F, and O labels describe typical, well-formed production messages. A means expected whenever the relevant message context applies, F means commonly present, and O includes optional, conditional, and implementation-specific fields. The labels are a practical reading guide, not a claim that every historical or malformed message contains the same fields.
| Category | Header or component | Description | Typical origin | Present |
|---|---|---|---|---|
| Identity | Date | Date and time the author or sending system composed the message. | Authoring client or service | A |
| Identity | From | Author identity presented in the message. | Authoring client or service | A |
| Identity | Sender | Agent that transmitted the message when it differs from the author, or when the From field contains multiple mailboxes. | Authoring client or service | O |
| Identity | Reply-To | Address where the author requests replies to be sent. | Authoring client or service | O |
| Identity | Organization | Organization associated with the author. This is not a core RFC 5322 field. | Authoring client or service | O |
| Identity | To | Primary recipients displayed in the message. This can differ from the SMTP recipient list. | Authoring client or service | F |
| Identity | Cc | Carbon-copy recipients displayed in the message. | Authoring client or service | O |
| Identity | Bcc | Blind-copy addressing information. It is normally removed or reduced to an empty field before delivery. | Authoring or submission system | O |
| Identity | Subject | Author-supplied message topic or summary. | Authoring client or service | F |
| Identity | Message-ID | Globally unique message identifier used for references, deduplication, and correlation. | Authoring client or submission service | F |
| Delivery | Return-Path | Finalized SMTP reverse-path used for delivery status handling. It is not the visible From address. | Final delivery system | F |
| Delivery | Received | Trace record prepended when a server accepts or relays the message. | Receiving or relaying server | A |
| Delivery | from, by, with, id, for, date | Clauses inside one Received field describing the sending host, receiving host, protocol, queue identifier, recipient, and timestamp. They are not separate header fields. | Receiving or relaying server | F |
| Delivery | Delivered-To | Mailbox or address selected during local delivery. Behavior varies by provider. | Recipient system | O |
| Delivery | User-Agent or X-Mailer | Software that composed or submitted the message. | Authoring client or service | O |
| Delivery | Return-Receipt-To | Nonstandard request for a receipt to be sent to an address. | Authoring client or service | O |
| Delivery | Disposition-Notification-To | Requests a message disposition notification. A request does not prove that a receipt will be generated. | Authoring client or service | O |
| Thread | In-Reply-To | Message-ID value of the message being answered. | Authoring client or service | O |
| Thread | References | Ordered Message-ID values that connect the message to its conversation history. | Authoring client or service | O |
| Resent | Resent-Date | Date and time a user or system reintroduced the message into transport. Required when a resent block is used. | Forwarding client or service | O |
| Resent | Resent-From | Identity of the person or system that resent the message. Required when a resent block is used. | Forwarding client or service | O |
| Resent | Resent-Sender | Transmitting agent when it differs from Resent-From, or when Resent-From has multiple mailboxes. | Forwarding client or service | O |
| Resent | Resent-To, Resent-Cc, Resent-Bcc | Recipients used for the resent message. | Forwarding client or service | O |
| Resent | Resent-Message-ID | Identifier assigned to this resent transmission. | Forwarding client or service | O |
| Resent | Resent-Subject | Nonstandard field sometimes created by software. It is not part of the RFC 5322 resent field set. | Forwarding client or service | O |
| MIME | MIME-Version | Declares MIME conformance. The defined value remains 1.0. | Authoring client or service | F |
| MIME | Content-Type | Specifies the media type, subtype, and parameters for the body or body part. | Authoring client or service | F |
| MIME | boundary parameter | Separator token required to delimit child parts in a multipart body. | Authoring client or service | O |
| MIME | Content-Transfer-Encoding | Specifies how one body part was represented for transport. It defaults to 7bit when omitted. | Authoring client or service | F |
| MIME | Content-Disposition | Suggests inline or attachment presentation and can supply a filename. | Authoring client or service | O |
| MIME | Content-ID | Identifier used to reference one body part from another, commonly through a cid URL. | Authoring client or service | O |
| Authentication | DKIM-Signature | Cryptographic signature containing the signing domain, selector, signed fields, body hash, and signature data. | Sending or signing system | F |
| Authentication | Authentication-Results | Machine-readable container for receiver results such as DKIM, SPF, DMARC, ARC, and BIMI. | Receiving authentication service | F |
| Authentication | Authentication-Results: dkim= | Reports whether one or more DKIM signatures verified and identifies the evaluated signing domain. | Receiving authentication service | F |
| Authentication | Received-SPF | Detailed receiver record of an SPF evaluation and the SMTP identity evaluated. | Receiving authentication service | O |
| Authentication | Authentication-Results: spf= | Reports the SPF result for the envelope sender or HELO identity. SPF policy itself is published in DNS. | Receiving authentication service | F |
| Authentication | Authentication-Results: dmarc= | Reports the DMARC result for the visible From domain after alignment is evaluated. DMARC policy itself is published in DNS. | Receiving authentication service | F |
| Authentication | BIMI-Selector | Selects a non-default BIMI DNS record for a message stream. The default BIMI record does not require this header. | Sending or authoring service | O |
| Authentication | Authentication-Results: bimi= | Reports the receiver BIMI evaluation and selected assertion record when the receiver implements BIMI. | Supporting receiving service | O |
| Authentication | BIMI-Location, BIMI-Indicator | Receiver-added fields that pass the validated indicator location or encoded SVG indicator to a trusted mail client. | Supporting receiving service | O |
| Authentication | ARC-Seal, ARC-Message-Signature, ARC-Authentication-Results | Authenticated Received Chain fields used to preserve handling and authentication evidence across intermediaries. | Participating intermediary | O |
DKIM can create a DKIM-Signature field in the outgoing message. SPF and DMARC do not create equivalent sender headers. A receiver normally records their results inside Authentication-Results, while SPF can also use Received-SPF. BIMI policy and indicator locations are discovered through DNS. The optional BIMI-Selector field can select a non-default record, and a supporting receiver can add BIMI result, location, and indicator fields for its trusted mail client.
How to read the delivery trace
Each server that accepts a message normally prepends a Received field. The newest trusted hop therefore appears near the top, while the earliest trace normally appears near the bottom. Read the chain from bottom to top when reconstructing the route, but do not assume every field is trustworthy. A sender can insert fake trace text before the first server under receiver control.
Compare the from, by, with, queue identifier, recipient clause, timestamp, and time zone at each trusted hop. Large time gaps can reveal queueing or processing delay. Hostname and address changes can reveal an unexpected relay, gateway, or route. Correlate the trace with MTA logs and SMTP responses whenever those records are available.
Why MIME is needed
The basic message format does not by itself describe HTML, attachments, alternative representations, or arbitrary binary files. The Multipurpose Internet Mail Extensions defined across RFC 2045 and related specifications add media types, character sets, transfer encodings, multipart containers, and body-part headers.
A MIME message can contain one body part or a tree of nested body parts. Every part has its own small header section, an empty line, and its own content. Multipart boundaries mark where child parts begin and end.
Single-part messages
A simple message can use one text/plain body:
From: [email protected]
To: [email protected]
Date: Mon, 19 Aug 2024 10:00:00 +0000
Message-ID: <[email protected]>
MIME-Version: 1.0
Content-Type: text/plain; charset="UTF-8"
Content-Transfer-Encoding: quoted-printable
Queue depth is back within the normal operating range.
The charset parameter explains how text bytes map to characters. The transfer encoding explains how the content was represented for transport. A parser must apply both correctly before displaying or indexing the text.
Multipart messages and nesting
A multipart body declares a boundary token in its Content-Type. Each child begins with two hyphens followed by that token. The final delimiter adds two more hyphens to close the container. Boundary values must not occur inside the enclosed content.
A common production message uses multipart/mixed at the outer level for the body plus attachments, then nests multipart/alternative for the plain-text and HTML versions:
From: [email protected]
To: [email protected]
Subject: Weekly delivery report
MIME-Version: 1.0
Content-Type: multipart/mixed; boundary="outer"
--outer
Content-Type: multipart/alternative; boundary="body"
--body
Content-Type: text/plain; charset="UTF-8"
Content-Transfer-Encoding: quoted-printable
The weekly report is attached.
--body
Content-Type: text/html; charset="UTF-8"
Content-Transfer-Encoding: quoted-printable
<p>The weekly report is attached.</p>
--body--
--outer
Content-Type: application/pdf; name="weekly-report.pdf"
Content-Transfer-Encoding: base64
Content-Disposition: attachment; filename="weekly-report.pdf"
JVBERi0xLjQK...
--outer--
The nesting is significant. Plain text and HTML are alternative representations of the same body, while the PDF is a separate sibling attachment. Flattening these parts can make a client display both alternatives, lose the attachment relationship, or misidentify the primary content.
Common multipart subtypes
RFC 2046 defines the core MIME media types and multipart behavior.
| Subtype | Purpose | Typical use |
|---|---|---|
multipart/mixed | Groups independent parts that should be kept together. | A message body followed by one or more attachments. |
multipart/alternative | Provides different representations of the same content in increasing order of preference. | Plain-text and HTML versions of one message. |
multipart/related | Groups a root document with resources needed to render it. | HTML plus inline images referenced by content identifiers. |
multipart/report | Carries a human-readable explanation and machine-readable report data. | Delivery status notifications and feedback reports. |
multipart/digest | Groups complete or encapsulated messages. | A collection of forwarded messages. |
multipart/signed | Pairs content with a detached signature. | S/MIME or OpenPGP signed mail. |
multipart/encrypted | Pairs control information with encrypted content. | OpenPGP encrypted mail. |
Media type, disposition, and filename
Content-Type describes what a body part contains. Common values include text/plain, text/html, image/png, image/jpeg, application/pdf, and message/rfc822. The type can include parameters such as a character set, boundary, name, or format.
RFC 2183 defines Content-Disposition. An inline part is intended for automatic presentation with the message. An attachment part normally requires a separate user action. The filename is a suggestion from an untrusted sender, so receiving software must sanitize paths, reserved names, control characters, and executable content before saving it.
An HTML body can refer to an inline resource through a content identifier:
Content-Type: image/png
Content-Transfer-Encoding: base64
Content-ID: <[email protected]>
Content-Disposition: inline; filename="chart.png"
The HTML can use cid:[email protected] as the image source. In a well-formed structure, the HTML and referenced resource normally share a multipart/related container.
Transfer encoding is not encryption
7bit: Content already fits the traditional line and character restrictions.quoted-printable: Mostly readable text is retained while unsafe bytes and line endings are encoded.base64: Binary content is represented with a transport-safe text alphabet. It is encoding, not secrecy or access control.8bitandbinary: These require compatible transport capabilities and must not be assumed across every route.
Decode one MIME part according to its own Content-Transfer-Encoding. Do not apply a parent encoding to every descendant, and do not guess from a filename extension.
Internationalized headers and addresses
Traditional header syntax is based on ASCII. Encoded-word syntax allows non-ASCII display text in fields such as Subject and display names. RFC 6532 also extends the message format to permit UTF-8 directly in header field values when the transport supports the internationalized email framework. Header field names themselves remain ASCII.
A parser must retain the original bytes, interpret the declared syntax, and avoid destructive normalization. Internationalized addresses and visually similar Unicode characters also require careful comparison in security and identity workflows.
Why preserving the raw message matters
For incident review, compliance, or archiving, preserve the complete raw message whenever policy allows. Keep the original header fields, order, folding, line endings, MIME boundaries, encoded body parts, and attachment metadata. A rendered HTML copy or screenshot loses evidence needed to reconstruct delivery and validate content handling.
Decoding a copy is useful for search and presentation, but it should not replace the source object. Content changes can also affect signatures. DKIM, S/MIME, and OpenPGP validation all depend on preserving the signed representation closely enough for the relevant canonicalization and verification process.
A practical message-analysis workflow
- Export the message as raw source or an
.emlfile instead of copying only the visible content. - Record the SMTP envelope, queue identifier, and server logs separately when they are available.
- Locate the empty line that separates the top-level header section from the body.
- Read trusted
Receivedfields from the earliest trusted hop toward the destination. - Identify the root
Content-Type, then follow each boundary recursively. - Decode every leaf part using its own transfer encoding and character set.
- Classify parts as body alternatives, related inline resources, attachments, reports, or encapsulated messages.
- Compare the visible author with the envelope sender, DKIM signing domain, authentication results, and final receiver policy.
- Retain the raw source and document every transformation made for display, indexing, or extraction.
Mistakes that make the message harder to read
- Treating the visible
Fromaddress as the SMTP envelope sender. - Expecting a delivered message to expose its Bcc recipient list.
- Reading the newest trusted
Receivedfield as the first hop. - Parsing multipart content without respecting nested boundaries and closing delimiters.
- Assuming HTML and plain text are independent messages instead of alternatives.
- Treating Base64 as encryption or as proof that an attachment is safe.
- Trusting a sender-supplied attachment filename as a local storage path.
- Saving only a rendered message when raw headers and encoded parts are needed for investigation.
What to keep from a message investigation
The top-level rule is simple: header fields, one empty line, then the body. MIME can turn that body into a tree of alternatives, inline resources, reports, complete messages, and attachments. SMTP carries the object with a separate transport envelope and adds delivery evidence around it.
Keep the untouched source. Work from a copy when unfolding fields or decoding parts. Follow trusted Received fields, preserve the boundary hierarchy, and compare transport identity with the visible author. That gives the next operator something a screenshot cannot: a message path that can be checked again.


