outtest
Get started
Channels

How to A/B test email subject lines

Send each subject line to a random slice of your list and pick the winner on clicks or revenue per recipient, not opens, which Apple Mail inflates since 2021.

Updated 28 September 2026 · 6 min read

The short answer

Split a random sample of your list, send each subject line to one group, and pick the winner on click rate or revenue per recipient, not open rate. Apple's Mail Privacy Protection, launched in 2021, loads tracking pixels whether or not someone reads the email, so opens are inflated and can crown the wrong subject line. To detect a 20% lift on a 2.5% click rate you need about 15,600 recipients per version, so smaller lists should split the whole send 50/50 and learn for next time.

Free tool: A/B test sample size calculator. No signup.

To A/B test email subject lines, send each subject line to a random slice of your list, then pick the winner on click rate or revenue per recipient. Don't pick it on open rate. Since Apple launched Mail Privacy Protection in 2021, Apple Mail loads the tracking pixel that counts opens for users who turn it on, even when nobody reads the email. Machine opens pad both versions, which buries real differences in noise, and a subject line that earns curious opens can still earn fewer clicks. The "winner" on opens is often the one that sells less.

Subject line tests also need more recipients than most people expect. At a 2.5% click rate, spotting a 20% lift takes about 15,600 recipients per version. If your list is smaller than about 30,000, split the whole send 50/50, judge on clicks and revenue, and carry what you learn into the next email.

Why open rate stopped working

Email tools count an open when a tiny invisible image in the email loads. Mailchimp describes it as "a tiny, transparent image" that it counts each time it's downloaded (about open and click rates).

In June 2021, Apple announced that in the Mail app, Mail Privacy Protection "stops senders from using invisible pixels to collect information about the user" and helps users "prevent senders from knowing when they open an email" (Apple Newsroom). In practice, Apple Mail preloads the pixel. Mailchimp's explanation is that Apple Mail "will preload pixels, even if your contact hasn't opened the email, resulting in unreliable open metrics", and it states that A/B test results based on open rates "may not be accurate" (Mailchimp MPP FAQ). Klaviyo also warns that open rates "may be inflated" (Klaviyo A/B testing help).

Clicks still work. Apple's own support page notes that Mail Privacy Protection "does not extend to links" (Apple Support).

What to judge a subject line on

Metric Still reliable? Use it for
Open rate No, inflated by Apple Mail preloading Spotting big problems only
Click-to-open rate No, its denominator is inflated opens Nothing, in tests
Click rate (clicks ÷ delivered) Yes, after bot filtering Deciding most subject line tests
Orders or revenue per recipient Yes Deciding tests on sends that sell
Unsubscribe and spam complaint rate Yes Guardrails

Clicks have their own noise. Security scanners and link preview tools sometimes click links before a person does, and Mailchimp says these bots "can falsely inflate open and click metrics". Turn on bot filtering if your tool has it (Mailchimp bot activity).

Why does a subject line affect clicks at all? Because people who never open can't click. The subject line changes who opens and what they expect to find, and that flows through to clicks and orders. Measuring at the click or order level captures the effect that matters.

How many recipients you need

The quick formula for a yes-or-no metric like click rate, at 95% confidence and 80% power, is n = 16 × p(1 − p) ÷ Δ² per version, where p is your usual click rate and Δ is the absolute difference you want to detect. It comes from Kohavi and colleagues' practical guide to controlled experiments.

Worked example

Your list has 40,000 subscribers and your click rate is 2.5%. You want to detect a 20% relative lift, from 2.5% to 3.0%, so Δ = 0.005.

  • n = 16 × 0.025 × 0.975 ÷ 0.005² = 0.39 ÷ 0.000025 = 15,600 per version
  • Two versions need 31,200 recipients, 78% of the list.

Now suppose you use the common setup of sending each version to 10% of the list (4,000 each) and the winner to the other 80%. Flip the formula to find the smallest lift that test can reliably see: Δ = 4 × √(0.025 × 0.975 ÷ 4,000) = 0.0099. That's about 1 percentage point, a 40% relative lift. Any smaller difference is likely to be noise, and the "winner" sent to 32,000 people is often a coin toss.

With a 40,000 list, the better plan is a 50/50 split of the whole send. Here is how that might look:

Subject A Subject B
Recipients 20,000 20,000
Open rate (inflated) 48% 41%
Clicks 420 (2.1%) 520 (2.6%)
Revenue $4,200 $5,800
Revenue per recipient $0.21 $0.29

On opens, A wins clearly. On clicks, B leads by 0.5 points. The standard error of that gap is √(0.021 × 0.979 ÷ 20,000 + 0.026 × 0.974 ÷ 20,000) = 0.0015, so z = 0.005 ÷ 0.0015 = 3.3, which is well past the usual 95% bar. Revenue per recipient agrees. B is the better subject line, and an open-rate test would have picked A. Use the sample size calculator with your own list size and click rate before you choose a split.

Setting up the test in your email tool

Most tools let you choose what decides the winner. Choose it yourself rather than accepting whatever is preselected.

  • Mailchimp tests subject line, from name, content or send time, with up to 3 variations. You can pick the winner by open rate, click rate, total revenue, or manually. The test must go to at least 10% of recipients, Mailchimp recommends at least 5,000 contacts per combination, and it suggests waiting at least 4 hours before sending the winner (create an A/B test).
  • Klaviyo offers open rate, click rate or placed order rate as the winning metric. Its help page still recommends open rate for subject line tests, while noting MPP inflation (Klaviyo A/B testing help). Choose click rate or placed order rate instead.

If you can't get enough recipients for an automatic winner, choose manual selection or a full 50/50 split, and review the result against your guardrails before the next send.

What to test in a subject line

Test ideas that could plausibly change who clicks, not tiny rewordings. Some pairs worth trying:

  • The offer stated plainly ("20% off all boots until Sunday") against a curiosity line.
  • A specific number or detail against a general promise.
  • The product or topic name first against the benefit first.
  • A short subject line against a longer one that previews the content.
  • Subject line plus preview text as a pair, since both show in the inbox.

Keep everything else identical: same send time, same from name, same content. If you change two things, you won't know which one worked. For a bigger list of ideas across channels, see what to A/B test first.

Step by step

  1. Decide the winning metric before you write anything: clicks, or revenue per recipient if your store is connected.
  2. Work out recipients per version with the formula or the calculator.
  3. Write two subject lines that differ in one idea.
  4. Randomize. Let the tool split the list, and don't send version A to one segment and B to another.
  5. Send both at the same time.
  6. Wait at least 4 hours for clicks, longer for orders.
  7. Check unsubscribes and spam complaints before rolling out the winner.
  8. Log the result, including losing tests, so the next test starts from what you learned.

Common mistakes

  • Letting the tool auto-pick on open rate.
  • Judging a 10% test slice with 1,000 recipients and calling a 0.3-point click gap a win.
  • Testing on your most engaged segment and applying the result to the whole list.
  • Sending the versions at different times, so the time of day decides the result.
  • Stopping early because one version leads after an hour, which is the peeking problem in email form.

Questions people ask

Are open rates useless for subject line tests?+

Mostly, yes. Apple's Mail Privacy Protection preloads the tracking pixel, so Apple Mail users can show as opened without reading anything. Mailchimp's own help pages say A/B test results based on open rates may not be accurate and suggest click rates instead. Opens can still hint at a problem, such as a sudden drop, but shouldn't decide a winner.

What should I use instead of open rate?+

Use click rate when you have a few thousand recipients per version, and revenue or orders per recipient when your email tool is connected to your store. Track unsubscribes and spam complaints as guardrails, so a subject line that wins on clicks by annoying people doesn't get rolled out.

How many recipients do I need per subject line?+

It depends on your click rate and the lift you want to detect. At a 2.5% click rate, detecting a 20% relative lift takes about 15,600 recipients per version at 95% confidence and 80% power. Mailchimp recommends at least 5,000 per combination as a floor.

How long should I wait before picking the winner?+

Mailchimp recommends at least 4 hours before sending the winner to the rest of the list. Clicks arrive faster than purchases, so if you judge on revenue, wait longer or treat the test as learning for the next send rather than an auto-pick.

Read next

Let Outtest run your split tests

AI agents read your analytics and payments, find where you lose the most money, build the fix and test it. Every test is judged on revenue, not clicks. Plans from $29 a month.