Email tests: what works best: Test one change at a time to isolate impact; Use tracked clicks as primary outcome for reliable results; Apply random assignment and control send windows
Image: Email Growth Desk

Email Copy

Subject line and email content experiments

Plan fair subject-line and email-content tests, choose useful outcomes and interpret results without overstating a winner.

A useful email experiment compares credible versions of a message sent to the same eligible audience. Before sending, decide what will change, which result matters and when you will assess it. A platform’s reported winner describes that test; it does not establish a rule for future campaigns.

Start with a decision the team can use

Frame the question as a choice: “Should this newsletter lead with the practical guide or the event invitation?” Both versions should be accurate and useful. If one is unclear or less relevant, improve it before testing.

A subject-line test compares ways to describe the same email. A content test compares information inside it. Changing the subject, offer, body and destination together may change the result, but it will not show which change mattered.

QuestionChangeKeep steadyPossible primary outcome
Which subject better sets up this message?Subject lineEligible audience, body, preview text, sender and send windowUnique people clicking the intended link
Which opening helps readers find the guide?Body openingSubject, audience, guide, destination and send windowUnique people clicking the guide link
Which explanation leads to an enquiry?One content sectionOffer, form, audience and send windowCompleted enquiries, if attributable to a variant

Choose an outcome that reflects the campaign’s purpose and can be measured credibly. A tracked link click does not establish that the destination loaded or that a later action occurred.

Make the comparison fair

Set the eligibility rule first. Apply permission and suppression rules, then randomly divide eligible recipients between versions. Check that a person appears in only one group.

Send both versions in the same window where practical. Record delivery or tracking problems that affect either group.

If only a small share of eligible recipients is assigned to the test, there may be too few in each group to learn much.

Control other influences where practical

A test can be affected by factors besides the version being compared. NIST describes these as nuisance factors and gives time of day as an example. Decide which such factors matter enough to track or control, rather than assuming random assignment removes every source of variation.

When a factor can be controlled, NIST describes blocking: form comparable groups in which that factor is held constant, then compare the test versions within each group.

For an email experiment, that could mean comparing both versions within each planned send-time group if timing cannot be kept the same. NIST’s general rule is to block what can be controlled and randomise what cannot.

Choose a result that matches the change

Opens are an imperfect measure of attention. Apple Mail Privacy Protection can preload tracking images, while blocked images can leave genuine reading unrecorded.

For a subject test, a click on the intended link may better indicate whether the subject brought people to the next step. For a body-copy test, an open cannot show whether the changed content helped.

Choose one primary metric, an observation window and a denominator in advance. Keep the numbers assigned to each version, successful deliveries, bounces, unsubscribes and complaints alongside the outcome.

Use enquiries or purchases only when they can be connected reliably to the assigned variant. A click towards either is not proof of completion.

Measuring Success: Primary Metrics by Test Type

Subject Line Test
Unique clicks on intended link
Body Copy Test
Clicks on key content links (e.g., guide, form)
Offer or Form Test
Completed enquiries or purchases (if trackable)

Read the result with its limits

Look at counts as well as rates. A difference of a few actions can be unstable. Automated security systems can also affect click figures, so check the available filtering and tracking settings. Do not replace the planned metric with whichever result looks most favourable after sending.

The decision might be use, repeat, revise or inconclusive. A credible result may support a similar campaign; a small or noisy difference may warrant another relevant test. If both versions disappoint, examine audience fit, offer and destination before testing more wording.

Record the question, exact variants, audience rule, send window, primary metric, raw results, measurement limits and decision. Related guides address audience control, metric choice, small-test interpretation and the learning record in more detail.

Add context to a proportion

When the primary outcome is a proportion, such as the share of delivered messages that produced the intended action, a confidence interval can put the observed rate in context. NIST describes the Wilson method for intervals on proportions and notes that its procedure does not strongly depend on the values of the proportion or sample size.

NIST also reports that Agresti and Coull recommended the Wilson method for virtually all combinations of sample size and proportion. An interval does not remove the need to inspect the underlying counts or consider measurement limits; it is another way to express the uncertainty around a measured proportion.

In this guide

  1. Testing a subject line without changing the audienceHold eligibility, timing and email content steady when comparing subject lines, then report the result with its limits.
  2. Choosing a useful success metric for an email testSelect a primary email-test outcome, define its numerator and denominator, and keep tracking limitations close to the decision.
  3. Avoiding misleading winners in small email testsRead small email-test results with raw counts, a fixed metric and uncertainty before declaring a variant the winner.
  4. Recording campaign learnings beyond the winning variantCapture the question, variants, audience, raw outcomes and limits of an email test so the next campaign can use the learning.

More from Email Copy