Email Campaign Performance Analysis: A Diagnostic Scorecard
Campaign analysis should explain what happened at each boundary and what changed because the email was sent. A single delivered, open or click rate cannot do that. Build a recipient-level funnel from eligibility through transport, safety, qualified action and business outcome, preserve denominators and terminal timing, then use randomized holdouts where practical to estimate incremental value.
Define a recipient-level funnel
eligible -> selected -> attempted -> accepted
-> deferred / rejected / expired
accepted -> qualified action -> conversion -> retained valueCount unique recipients at each boundary and keep retries attached to the same recipient. State whether suppressed rows are outside eligible or reported as exclusions. An SMTP 2xx proves receiver acceptance, not inbox placement. Keep queued recipients unresolved until accepted or expired.
Publish formulas and denominators
| Metric | Example definition |
|---|---|
| Acceptance rate | Accepted recipients / attempted recipients |
| Complaint rate | Provider or program complaints / documented denominator |
| Qualified click rate | Recipients with filtered qualified click / accepted recipients |
| Conversion rate | Converting eligible recipients / eligible recipients |
| Incremental lift | Treatment outcome minus comparable holdout outcome |
Label windows, deduplication and scanner filtering. Do not compare percentages with different denominators as if equivalent.
Use opens only with explicit limitations
Image requests can be blocked, cached or prefetched by privacy services. They cannot reliably prove a human read. Keep opens as diagnostic data with client/proxy classification, but avoid using them as the primary success metric or as the sole inactivity rule.
Clicks also need bot and security-scanner controls. A qualified action may require a browser session, plausible sequence, authenticated event or downstream behavior. Version the filter and apply it consistently across variants.
Segment where an action can follow
Break results down by receiving organization, acquisition cohort, lifecycle state, campaign version, sending identity and device/client when available. First compare selection and transport so a creative conclusion is not drawn from a provider rejection. Then compare safety and business outcomes.
Avoid slicing until a random high appears. Predefine important segments and apply minimum sample and uncertainty. Small regional providers can be operationally important even when global totals look stable.
Use holdouts and clean randomization
Randomize at the recipient or account level before treatment, prevent cross-variant contamination and preselect a primary outcome. Check balance in provider, lifecycle and recent activity. Preserve assignment even when a recipient is not delivered, then report both intention-to-treat and appropriate operational diagnostics.
For revenue, define attribution and return/cancellation handling. Incremental margin can be more useful than attributed gross revenue, especially with discounts.
Use a boundary-first diagnostic sequence
- Did eligible and selected counts match the rule?
- Was the intended production artifact attempted?
- Did acceptance, deferral or rejection change by provider?
- Did complaint and unsubscribe guardrails change?
- Did qualified actions change after filtering?
- Did conversion and incremental value change?
- Were tracking, product or warehouse data complete?
This order prevents a broken redirect from being called poor content and a provider deferral from being called weak demand.
Define the campaign population and event contract first
campaign_assignment(
campaign_id, subject_key, variant, holdout,
assigned_at_utc, eligibility_version
)
recipient_delivery(
campaign_id, subject_key, provider_org,
attempted_at, final_smtp_class, enhanced_reply,
outbound_ip, message_id
)
outcome(
subject_key, outcome_id, occurred_at_utc,
type, gross_value, discount, refund, variable_cost
)Keep selected, excluded, assigned, attempted, accepted, deferred, rejected and expired populations. Store event time separately from ingestion. Deduplicate on stable IDs and retain metric/classifier versions.
Separate transport, safety, interaction and business measures
| Layer | Metric examples | Decision |
|---|---|---|
| Transport | Acceptance, deferral, rejection, queue age | Can receiver accept? |
| Safety | Complaint, unsubscribe, hard bounce | Is audience/program healthy? |
| Interaction | Qualified click, reply, account action | Observable action? |
| Business | Conversion, retention, net margin | Desired outcome? |
| Causal | Lift versus randomized holdout | What did campaign create? |
Do not combine these into one score or call SMTP acceptance inbox placement.
Publish numerator, denominator, window and unknowns
mx_acceptance_rate = accepted_recipients / attempted_recipients
complaint_rate_internal = complaints / chosen_internal_denominator
qualified_click_rate = subjects_with_qualified_click / assigned_or_accepted
conversion_rate = matured_converters / assigned_subjects
provider spam rates may use provider-defined denominators;
do not force equality with ESP metricsOne recipient can generate several retry events. Use unique recipient/assignment denominators for people and event counts for workload. Show sample size and data delay beside percentages.
Classify privacy fetches and security-scanner clicks
Apple Mail remote-content fetches do not prove a human read. Gateways can click links before delivery. Preserve raw events and assign supported, machine-likely or unknown through a versioned classifier. Do not silently delete or label remaining events as human.
| Event | Analysis use |
|---|---|
| Remote image fetch | Weak diagnostic/client behavior |
| Known scanner click | Security-delivery clue, exclude from human intent |
| Browser-correlated click | Qualified interaction under definition |
| Authenticated action | Stronger downstream behavior |
Annotate trends when classifier/client mix changes.
Keep attributed credit separate from incrementality
First-touch, last-touch and multi-touch assign conversion credit under a model; they do not show what would happen without email. Reconcile order IDs, refunds, discounts, tax treatment and variable cost to finance. Avoid adding full revenue reported by several platforms.
absolute_lift = treatment_rate - control_rate
incremental_outcomes = absolute_lift * treatment_assigned
incremental_margin = treatment_net_margin
- expected_control_net_margin
Preserve intention-to-treat and report uncertainty. A high-intent cohort can generate large attributed revenue with little causal lift.
Analyze providers, cohorts and sources without post-treatment bias
Preplanned breakdowns include receiving organization, acquisition source, lifecycle, permission age, variant and eligible treatment. Avoid declaring a result from dozens of uncorrected subgroups. Clickers versus non-clickers is not a causal comparison because clicking happens after treatment.
| Breakdown | Question |
|---|---|
| Provider | Transport or complaint concentration? |
| Source | Permission/quality difference? |
| Variant | Randomized treatment difference? |
| Lifecycle | Heterogeneous value/risk? |
| Route | Infrastructure comparison if assignment supports it? |
Reconcile the campaign pipeline before reading performance
- Selected equals eligible plus documented exclusions.
- Assigned variants/holdout match randomization plan.
- Attempted plus final pre-send cancellations match export/ESP acknowledgment.
- SMTP final states reconcile unique recipients.
- Interactions map to valid message/subject identities.
- Conversions deduplicate and mature through refund window.
Alert on missing provider, unknown reply families, unmatched conversions, duplicate orders and event-ingestion lag. A beautiful dashboard cannot repair a broken denominator.
Worked case: open rate rises while qualified outcomes fall
A redesigned newsletter reports opens increasing from 28% to 47%, but qualified clicks and conversion decline. Client mix shifted toward Apple Mail and the new template loads more remote images. The team does not call the campaign a success.
It checks SMTP acceptance and rendering, versions proxy-fetch classification, compares qualified interactions and uses a randomized holdout for incremental orders. The redesign has no incremental lift and slightly higher complaints among one acquisition source.
The source and template are corrected, and the report keeps raw fetch rate as a diagnostic. The decision follows safety and business evidence, not the visually largest metric.
Build a campaign decision dashboard
- Eligibility, assignment, exclusions and holdout integrity.
- Provider-level attempted/accepted/deferred/rejected and queue age.
- Complaints, unsubscribes and hard bounces with counts.
- Qualified interaction/classifier version and unknowns.
- Matured conversion, refunds, cost and net margin.
- Attributed credit separately from incremental lift.
- Source/lifecycle/variant breakdowns under planned analysis.
- Release, data-quality and definition annotations.
Campaign analysis checklist
- Name the decision and primary outcome before launch.
- Keep immutable assignment and event contracts.
- Publish denominators and maturity windows.
- Separate transport, placement, interaction and value.
- Classify privacy/scanner noise.
- Reconcile finance and deduplicate orders.
- Use holdouts for causal claims.
- Report uncertainty and negative outcomes.
- Version metrics and preserve raw events.
Design A/B tests with power, balance and one primary decision
Define randomization unit, population, variants, primary outcome, guardrails, maturity and analysis before send. Estimate sample size from baseline, minimum detectable effect, significance and power; do not use a universal “1,000 per variant” rule.
rough two-proportion planning inputs:
baseline conversion p0
minimum meaningful absolute difference delta
alpha (false-positive control)
power 1-beta
use a validated statistical tool/library;
retain assumptions and achieved sampleCheck sample-ratio mismatch before interpreting outcomes. Repeated peeking and multiple metrics inflate false positives.
Investigate sample-ratio mismatch before reading lift
If assignment expected 50/50 but observed eligible/delivered groups differ materially, inspect randomizer, filtering, identity duplication, platform suppression, provider routing and logging. Treatment-caused delivery failure remains an outcome, but missing assignment or pre-treatment imbalance can invalidate comparison.
| Boundary | Count |
|---|---|
| Assigned | Authoritative randomization population |
| Eligible after pre-existing gates | Should not depend on variant |
| Attempted/accepted | May reflect treatment/route effects |
| Matured outcome | Analyze by assignment |
Do not rebalance by deleting rows post hoc.
Analyze provider transport without calling it placement
Segment attempts, acceptance, deferrals, rejections and queue age by provider, route and variant. If variants use the same route but message size/content differs, delivery can be treatment-dependent and belongs in the result. Seeds can support a placement hypothesis but are not population rates.
provider_variant(
assigned, attempted, accepted,
temporary_deferred, permanent_rejected,
expired, complaint, qualified_outcome
)Preserve complete replies. Do not exclude a variant’s bounces/filtering from intention-to-treat to make conversion look better.
Reconcile outcomes to finance and late adjustments
Define gross versus net revenue, discount, tax, refunds, cancellations, cost of goods/fulfillment and currency conversion. Use unique order IDs and a close/restatement schedule. Analytics event time, order time and settlement time differ.
| Bridge | Evidence |
|---|---|
| Transaction coverage | Matched/unmatched IDs |
| Value | Gross, net, margin and refund |
| Time | Event/order/settlement/maturity |
| Identity | Matched, duplicate, shared account |
| Currency | Source and exchange-rate version |
Do not force analytics and finance totals to equality; explain differences.
Investigate a performance anomaly at the earliest changed boundary
- Confirm data freshness, definitions and pipeline counts.
- Check eligibility/assignment and population mix.
- Check submission, provider SMTP and queues.
- Validate rendering, links and scanner/privacy classification.
- Check conversion site/product availability.
- Reconcile finance and attribution maturity.
- Compare campaign/source/infrastructure releases.
Pause harmful traffic when safety changes; otherwise preserve the experiment. A tracking outage should not trigger an unnecessary IP/domain migration.
Compare campaigns with seasonality and mix visible
Campaign-over-campaign trends mix weekday, season, promotion, audience, provider and measurement changes. Use matched cohorts or experiments for decisions. Annotate client-classifier, attribution and source shifts. Report confidence ranges and historical restatements.
comparison dimensions:
eligibility_version, source_mix, lifecycle_mix,
provider_mix, weekday/season, offer,
route_identity, metric_version, maturity
A year-over-year chart is descriptive unless assumptions support causal inference.
Minimize recipient-level analytics data
Use stable pseudonymous keys, restricted access, purpose/retention and aggregate outputs. Do not upload raw customer/message data to unapproved tools. Respect deletion/consent changes in derived tables and experiment exports.
Scanner/IP/user-agent data can be personal or sensitive; collect only what a documented classifier/decision needs. Keep raw and derived access separate. Analysts should use controlled samples for incident investigation.
Report small cohorts carefully to prevent re-identification.
Write an executive conclusion that maps evidence to a decision
State population, treatment, primary outcome, lift/uncertainty, safety, transport, data quality and business implication. Example: “Among 180,000 assigned eligible accounts, variant B produced no detectable incremental margin; complaints rose at Yahoo in partner-source recipients; rollout is stopped while that source is corrected.”
Avoid “delivery 99%, opens up, campaign successful.” Name what delivery means, mark privacy limits and keep attributed revenue separate from incremental value. List decision owner and next evidence date.
Publish a versioned metric dictionary with every report
Each metric needs a business name, precise numerator, denominator, eligible event types, deduplication key, event and processing time, maturity window, classifier version, owner and known limitations. “Delivery rate” is ambiguous unless it states whether the denominator is selected, attempted or accepted and whether temporary deferrals have matured.
| Metric | Required qualification |
|---|---|
| Accepted | Remote SMTP acceptance, not inbox placement |
| Complaint rate | Provider feedback coverage and denominator |
| Qualified click | Bot/scanner classifier and uniqueness rule |
| Conversion | Event definition, identity and window |
| Revenue/margin | Refund, discount, cost, currency and maturity |
Definition changes create a new version and annotated trend break; they do not silently rewrite history.
Control repeated looks, many variants and many outcomes
Checking an ordinary fixed-horizon test every hour and stopping when significance appears increases false positives. Use the preplanned sample/maturity or a valid sequential method. If many subjects, metrics, segments or variants are tested, define a correction or hierarchy before seeing results.
Guardrails such as complaints and severe delivery failures can stop a harmful campaign early even when the primary business test is incomplete. That is an operational safety decision, not proof that another variant wins.
Report estimate, confidence interval, sample and practical threshold, not only a p-value. A statistically detectable lift can be too small to cover implementation, incentive and support cost; an inconclusive result is not evidence of equivalence.
Keep cohort maturity and late outcomes visible
A campaign sent over several days gives early recipients more time to convert, refund or complain. Build outcome windows from assignment or accepted time according to the question, and compare cohorts only after the same maturity. Preserve event-time and ingestion-time so late data can be restated.
cohort report columns:\nassigned_date, variant, assigned_count, accepted_count,\noutcome_window_days, matured_count, conversions, refunds,\nnet_margin, complaints, data_watermark, report_versionPublish provisional and closed results separately. Do not mix a seven-day treatment outcome with a one-day holdout outcome or freeze revenue before the normal refund window.
Audit identity matching before trusting attribution
Email address, browser cookie, login, customer, household and account represent different units. Shared devices, forwarded messages, cross-device sessions and identity merges can attach outcomes incorrectly. Define deterministic and modeled matches separately and show unmatched volume.
| Risk | Diagnostic |
|---|---|
| Duplicate order credit | Unique order ID across touches/campaigns |
| Shared account | Account versus person assignment |
| Cookie loss | Authenticated and unmatched outcome trend |
| Identity graph change | Version and before/after reconciliation |
| Forwarded message | Do not assume clicker equals recipient |
Experimental assignment at the decision unit remains more reliable for causal lift than expanding a modeled attribution graph after the campaign.
Make analytical queries reproducible and reviewable
Store the report code, parameters, data snapshots/watermarks and metric versions. Use tests for one-to-many joins, duplicate events, null identity, timezone boundaries and late corrections. Reconcile every join to counts before and after.
SELECT campaign_id, variant,\n COUNT(DISTINCT assigned_subject) AS assigned,\n COUNT(DISTINCT CASE WHEN qualified_conversion THEN order_id END) AS orders,\n SUM(CASE WHEN qualified_conversion THEN net_margin ELSE 0 END) AS net_margin\nFROM point_in_time_campaign_outcomes\nWHERE maturity_days >= :required_days\nGROUP BY campaign_id, variant;This illustrative query still requires assignment-unit and shared-order policy. Code review should reject a report that joins clicks to orders many-to-many or filters treatment delivery failures after randomization.
Final campaign-analysis signoff
- The decision, population, assignment unit and primary outcome were declared.
- Eligibility, exclusions and assignment reconcile.
- Provider SMTP data is separated from placement claims.
- Privacy fetches and automated clicks are classified and versioned.
- Outcome identity, attribution window and maturity are stated.
- Holdout contamination and sample-ratio mismatch were checked.
- Refunds, discounts, costs and currency reconcile to finance.
- Negative outcomes use counts and appropriate denominators.
- Uncertainty, multiple testing and data-quality limitations are visible.
- The conclusion names an action, owner and review date.
A useful report tells operators what changed and decision-makers what evidence supports the next action. A page of rates with undefined denominators does neither.
Run the campaign review from decisions and evidence, not dashboard order
Open with the planned business decision, assignment population, primary outcome, guardrails and maturity. State whether the result is final, provisional or invalidated by data quality. Show treatment effect with uncertainty and the minimum meaningful effect; then show incremental margin and recipient harm. This prevents a visually large open-rate change from controlling the meeting.
Next, walk the delivery boundary by provider: assigned, attempted, accepted, deferred, rejected, expired and complained. Use complete SMTP evidence and queue age. If placement is discussed, label seed, panel or provider evidence accurately and state coverage. Do not turn accepted mail into an inbox-rate claim.
Then inspect audience and execution dimensions chosen before analysis: acquisition source, lifecycle state, provider, variant and route. Surface sample-ratio mismatch, identity uncertainty, tracking/classifier changes, site incidents and finance reconciliation. Exploratory segments can generate a hypothesis but should not be presented as confirmed winners after searching many cuts.
Close with one of four decisions: scale, hold, change and retest, or stop. Assign owners for deliverability, data, product/site, creative and lifecycle defects. Record the next release ceiling and rollback trigger. Archive the report code, snapshots, definitions and decision. The meeting succeeds when operations know exactly what will change and which future evidence will confirm it.
Keep a decision ledger for recurring campaign analysis
Record the hypothesis, approved change, expected direction, affected population, release date, observation window and owner. At the next review, reconcile the promised change against what the sending system actually executed. This prevents repeated recommendations with no implementation evidence and distinguishes a strategy failure from a deployment that never occurred.
Link the ledger to report code, metric versions, maturity state, provider observations and rollback decisions. Retire dashboards that no longer map to an active decision while retaining governed history. When a result is restated because of late refunds, identity correction or classifier change, preserve both versions and explain why the decision did or did not change.
Archive the evidence needed to reproduce the conclusion
Retain governed campaign definitions, assignment extract, query version, data watermark, metric dictionary, finance close, provider evidence and decision record. Access should be restricted and retention appropriate. A future analyst must be able to reproduce the published totals without relying on an editable dashboard snapshot.


