Comparability methodology
Email Benchmark Comparability Worksheet
You can compare an external benchmark with your campaign only when its formula, population, measurement conditions, and source are close enough for the decision you need to make. Complete the worksheet first; then choose use, reject, or build your own baseline.
How to use this worksheet
Complete one record for one external benchmark and one campaign. Do not fill gaps with assumptions. Record every mismatch, name the decision the comparison would support, and choose one of the three outputs below.
Download the worksheet CSVUse
fields match the decision
Reject
a material mismatch remains
Build
your own dated baseline
Executive summary.
The short answer is conditional: an outside number is comparable only when you can reproduce what was counted, who and what was measured, where and when it was measured, and how collection or privacy controls changed the observation.
Capture the metric definition, numerator, denominator, audience, message type, mailbox-provider scope, geography, window and date, instrumentation and privacy effects, and source quality. If a missing or incompatible field could change the decision, reject the comparison instead of presenting it as an industry average.
When no external source passes, define the metric once and build a segmented first-party baseline. The downstream measurement workflow belongs in the Email Analytics Guide; this page only decides whether an external reference is fit for comparison.
Methods.
Start with a decision, not a number. Write the action the benchmark may influence, the campaign being assessed, and the maximum mismatch that would still leave that action defensible.
Transcribe the external source before looking at your result. Preserve the source's labels and unknowns, then complete the same fields for your campaign. This reduces the temptation to redefine a metric after seeing the gap.
Classify each field as matched, acceptably different with a written reason, materially different, or unknown. A materially different or unknown decision-critical field fails the comparison.
Return exactly one output: use the benchmark for the named decision, reject it, or build a first-party baseline. Record the source URL, access date, measurement window, decision owner, and review date with the output.
Download data
Complete one row for every comparison field. Use matched, acceptable, material, or unknown for status, then choose use, reject, or build as the final output.
Download the worksheet CSVStep 1 · Formula
Write the metric as a reproducible fraction.
A familiar label is not a definition. Record what qualifies for the numerator, what is eligible for the denominator, and how duplicates, retries, and missing events are handled.
Capture the metric definition
Write the exact event being measured and the unit of analysis: message, delivered message, recipient, account, session, or attributed outcome. Note whether the value is unique, total, gross, or net and whether automated activity is filtered.
Capture the numerator
Record the counted event, deduplication rule, qualification rule, and observation deadline. For replies or conversions, state whether every event counts or only a reviewed class such as positive replies or qualified outcomes.
Capture the denominator
Record whether the base is attempted, accepted, delivered, opened, clicked, or eligible recipients. For bounce measures, state how temporary 4.X.X and permanent 5.X.X SMTP outcomes are classified rather than treating every delivery status as equivalent.
Step 2 · Population
Match the people, messages, providers, and places.
A correct formula can still produce a misleading comparison when the underlying population or operating context differs.
Record audience and list context
Capture relationship, acquisition or consent path, segment, lifecycle stage, organization type, list age, and relevant sender maturity. A subscriber campaign, customer message, and unrequested prospecting campaign should not share a benchmark merely because each uses email.
Separate message type
Classify the message as prospecting, subscribed marketing, lifecycle, transactional, or relationship mail and preserve the source's own categories. The FTC says US CAN-SPAM covers commercial messages, including B2B, while it treats transactional or relationship categories narrowly; legal scope is not interchangeable with metric comparability.
Record mailbox-provider scope
Capture the destination-provider mix and whether the source covers all mailboxes or a provider-specific dataset. Gmail Postmaster data concerns mail to personal Gmail accounts, can omit low-volume data for privacy, and uses provider-defined dashboard calculations; Yahoo publishes its own sender practices. Do not relabel either source as a universal email population.
Sources: Email sender guidelines, Postmaster Tools dashboards, Sender best practices
Record geography and recipient type
Capture sender and recipient geographies, jurisdiction assumptions, and whether recipients are consumers, corporate subscribers, sole traders, or another class relevant to the comparison. The ICO distinguishes UK corporate subscribers from sole traders and says its B2B marketing guidance is under review after legal changes, so record the access date and obtain current legal review where needed.
Source: Business-to-business marketing
Record the window and date
Write the send-date range, event observation period, reporting time zone, lag, season, and publication or extraction date. Keep a point-in-time result separate from a rolling average and from a benchmark aggregated across undisclosed periods.
Source: Postmaster Tools dashboards
Test population compatibility
List each population difference and explain why it would or would not change the intended decision. If the source does not disclose a decision-critical population field, mark it unknown; a broad category label is not evidence of a match.
Step 3 · Measurement
Audit instrumentation and privacy effects.
Compare how each side observed the event, not only the values produced by the tools.
Treat opens as instrument-dependent
Record the tracking method, client mix, unique-event rule, bot or proxy treatment, and privacy features. Apple says Mail Privacy Protection prevents senders from seeing whether a protected recipient opened a message and hides the recipient's IP address, so open observations collected under different client and privacy mixes may not be comparable.
Define click and conversion attribution
Record redirect or analytics tooling, bot filtering, attribution model, lookback window, and UTM taxonomy. Google Analytics documents campaign parameters for identifying referral traffic and warns that missing relevant parameters can create unset reporting values.
Keep personal data out of campaign labels
Document privacy filtering and redaction before comparing attributed results. Google says Analytics data must not contain information it could recognize as personally identifiable and specifically warns against putting PII in custom campaign parameters.
Source: Best practices to avoid sending personally identifiable information
Record provider-dashboard limits
Note coverage, authentication prerequisites, latency, time zone, privacy suppression, and calculation scope. Gmail Postmaster says its data is not real time, some dashboards cover DKIM-authenticated mail, low volume can suppress data, and some dashboard datasets differ slightly.
Source: Postmaster Tools dashboards
Separate observed from inferred events
Label direct events, modeled or attributed events, provider feedback, and analyst classifications separately. Do not combine them into one rate unless the external source uses the same rules and the combination supports the named decision.
Test measurement compatibility
Explain whether instrumentation differences are immaterial, can be normalized, or invalidate the comparison. Keep any adjustment formula and the unadjusted values so another reviewer can reproduce the decision.
Step 4 · Evidence
Grade source quality without inventing precision.
A benchmark can be well documented yet unsuitable, or relevant yet too opaque to use. Judge both evidence quality and decision fit.
Record provenance
Save the original URL, publisher, named author or owner, publication and update dates, access date, and whether the source is a primary standard, provider or regulator document, first-party dataset, vendor analysis, or unsourced roundup.
Inspect the sample and exclusions
Look for sample construction, unit of analysis, inclusion and exclusion rules, deduplication, weighting, missing-data handling, and aggregation. Record each undisclosed field as unknown rather than estimating it.
Check reproducibility and conflicts
Ask whether another reviewer could calculate the result from the disclosed method. When sources conflict, preserve both definitions and investigate the scope difference; do not average unlike numbers to create false agreement.
Step 5 · Decision
Return use, reject, or build your own baseline.
The output is a documented decision for one use case, not a permanent quality score for the source.
Use
Use the benchmark only for the named decision when every decision-critical field is present and matched, or a difference is explicitly judged immaterial. Attach the completed record and describe the value as source- and scope-specific.
Reject
Reject the comparison when a formula, population, window, instrumentation, or source-quality mismatch could change the action, or when a critical field is unknown. Keep the source in the record so the same unsuitable comparison is not repeated.
Build your own baseline
When external evidence does not pass, freeze your metric definition, segment by the fields that matter, preserve campaign and measurement changes, and collect a dated first-party series. Report uncertainty and revisit segmentation when the operating context changes.
Practical checklist
Decision: name the action this comparison may change.
Metric definition: state the event, unit, uniqueness, and exclusions.
Numerator: record what counts, qualification, deduplication, and deadline.
Denominator: record the eligible base and delivery-status treatment.
Audience: record relationship, acquisition path, segment, stage, and list context.
Message type: separate prospecting, subscribed marketing, lifecycle, transactional, and relationship mail.
Provider scope: record the destination mix and provider-specific dataset limits.
Geography: record sender and recipient locations plus relevant recipient class.
Window and date: record send range, observation period, time zone, lag, season, and access date.
Instrumentation and privacy: record tracking, filtering, attribution, client mix, suppression, and redaction.
Source quality: record provenance, sample, exclusions, aggregation, missing data, and reproducibility.
Output: choose use, reject, or build your own baseline and assign an owner and review date.
Related resources
Source ledger
Material provider and compliance claims link to current primary sources beside the relevant section. This ledger records what each source supports.
Google · verified 2026-09-01
Provider-specific sender scope and the requirement to interpret Gmail evidence within Google's published sending context.
Google · verified 2026-09-01
Personal-Gmail scope, dashboard definitions, authentication coverage, UTC reporting, latency, privacy suppression, and dataset differences.
Apple · verified 2026-09-01
The effect of Mail Privacy Protection on sender visibility into opens, IP address, activity, and location.
Yahoo · verified 2026-09-01
Yahoo's provider-specific sender practices and the need to preserve destination-provider scope.
IANA · verified 2026-09-01
Standard success, persistent transient failure, and permanent failure classes used when defining delivery and bounce measures.
Google Analytics · verified 2026-09-01
UTM campaign identification, parameter meanings, traffic-acquisition reporting, and the reporting effect of missing relevant parameters.
Google Analytics · verified 2026-09-01
Analytics restrictions on personally identifiable information, including URLs and custom campaign parameters.
Federal Trade Commission · verified 2026-09-01
US commercial-email scope, B2B coverage, and the narrow treatment of transactional or relationship messages.
Information Commissioner's Office · verified 2026-09-01
UK B2B recipient classifications, electronic-marketing context, data-protection considerations, and the notice that this guidance is under review.
Maintenance owner
Folderly Research and Revenue Operations
Review quarterly and whenever a listed provider, analytics platform, or regulator changes the cited guidance. Next review: 2026-11-30.
Review triggers
- • A cited source changes its metric, scope, privacy, or compliance language.
- • The Email Analytics Guide changes its measurement definitions or attribution workflow.
- • A review finds that a worksheet field is routinely unknown or fails to change the output.
- • A new first-party data source changes which campaign segments can support an internal baseline.
Caveats and limits
This worksheet does not publish universal averages, performance promises, rankings, or a claim that one source predicts inbox placement, replies, or revenue.
Comparability is decision-specific. A source accepted for directional planning may still be unsuitable for forecasting, target setting, compensation, or causal claims.
Provider dashboards, email platforms, analytics tools, and privacy controls observe different events and populations. A clean-looking percentage does not remove those differences.
The compliance references are scoping inputs, not legal advice. Verify current law, recipient classification, and regulator guidance for the jurisdictions involved.
A first-party baseline can be more relevant than an external benchmark but can still be unstable, biased, or confounded. Preserve definition changes and disclose uncertainty.
Build the measurement workflow behind your baseline.
Once you have accepted a definition or chosen to build your own reference, use the Email Analytics Guide to set campaign parameters, collect the right events, segment results, and run a dated review.