outtest
Get started
What to test

How to A/B test headlines

Write headlines that make different promises, split traffic evenly and judge on revenue, not clicks. What to try, how long to run it and what can go wrong.

Updated 28 September 2026 · 5 min read

The short answer

Write two to four headlines that make genuinely different promises, split visitors evenly, and judge on revenue per visitor rather than clicks or signups. Run for at least one to two full weeks and until you reach the sample size for the lift you care about. Headlines can matter a lot: a change to how Bing displayed ad headlines raised its revenue 12%, which the researchers put at over $100 million a year in the US.

Free tool: A/B test significance calculator. No signup.

To A/B test headlines, write two to four headlines that make genuinely different promises, show each to an equal random share of visitors, and judge them on revenue per visitor, not on clicks or signups. Keep the rest of the page the same, or change only what the new promise needs, such as the subheadline. Run the test for at least one to two full weeks so every weekday is covered, and until each version has the sample size for the lift you care about.

The biggest mistake is testing word swaps. "Fast invoicing software" against "Quick invoicing software" will almost never produce a difference you can measure. "Get paid in half the time" against "Invoicing for plumbers and electricians" might, because they appeal to different reasons to buy.

Why headlines are worth testing

The headline is often the first thing a visitor reads, and it decides what they think the page is about. Small wording changes can move a lot of money at scale. In a 2017 Harvard Business Review article, Ron Kohavi and Stefan Thomke describe a Bing experiment that changed how ad headlines were displayed. It raised revenue by 12%, which they put at more than $100 million a year in the US. Before the test, the idea had been rated a low priority.

Publishers test headlines constantly. The Upworthy Research Archive is a public record of 32,487 headline and image experiments that Upworthy ran between January 2013 and April 2015, covering 150,817 versions and over 538 million visitor assignments. It is a useful place to see how differently worded headlines for the same story perform.

That data also shows the risk of optimizing for clicks. A 2023 study in Nature Human Behaviour analyzed about 105,000 Upworthy headline variations and found that, for a headline of average length, each additional negative word raised the click-through rate by 2.3%, while positive words lowered it. For a news site selling clicks, that is useful. For a business selling a product, a headline that wins clicks by alarming people can bring in visitors who don't buy.

Write different promises, not different words

Each headline in a test should give a different reason to care. Here are angles to try, with examples for a made-up invoicing tool:

Angle Example
Outcome "Get paid in half the time"
Pain "Stop chasing late invoices"
Audience "Invoicing for plumbers and electricians"
Specific number "Send an invoice in 60 seconds"
Ease "Invoices that write themselves from your jobs list"
Objection "Switch from spreadsheets in one afternoon"
Comparison "Everything your accountant wishes you used"

Only test claims you can back up. "Get paid in half the time" needs real data behind it. Unsupported claims can win a test, then cost you refunds, complaints and, in some markets, trouble with advertising rules.

Keep the page consistent with the promise. If the headline says "for plumbers", the screenshot, examples and reviews should show plumbers. If traffic comes from ads, the headline should match what the ad said.

A worked example

This is a made-up example. An invoicing tool with a $50 monthly plan tests three headlines, with 5,000 visitors each over two weeks. The page offers a free trial, and 45 days later the team checks who paid.

A: "Invoicing software for small businesses" B: "Get paid in half the time" C: "Invoicing for plumbers and electricians"
Visitors 5,000 5,000 5,000
Trial starts 150 (3.0%) 180 (3.6%) 140 (2.8%)
Paid within 45 days 45 (30% of trials) 45 (25%) 56 (40%)
Revenue $2,250 $2,250 $2,800
Revenue per visitor $0.45 $0.45 $0.56

On trial starts, B looks like the winner. The difference between 3.6% and 3.0% has a standard error of the square root of (0.03 × 0.97 ÷ 5,000 + 0.036 × 0.964 ÷ 5,000), about 0.36 percentage points. The gap is 0.6 points, which gives B about a 95% chance of having the higher trial rate. Yet B brought in exactly the same revenue as A. Its bigger promise attracted more people who didn't pay.

C won fewer trials and the most revenue, because it spoke to the people most likely to buy. Is it proven? Comparing paid customers, 56 out of 5,000 (1.12%) against 45 out of 5,000 (0.90%), gives C about an 86% chance of beating A. That is promising, not settled. The right move is to keep the test running until C clears your bar or stops looking good. Check your own numbers in the significance calculator.

How long to run a headline test

  • At least one full week, and two if your traffic varies a lot by day. Weekend visitors can behave differently from weekday visitors.
  • Until each version reaches the sample size you planned. The sample size guide shows how to set it.
  • Without stopping the moment one version looks ahead. Checking every day and stopping on the first good result inflates false wins. The peeking problem guide explains why.

Outtest sets these rules in advance for every test. By default a winner needs at least 7 days, a 90% chance of beating the original and a lift of at least 10% in revenue per visitor, and a test that hasn't cleared all three by day 42 is called a draw. Statistical significance explained covers the maths behind bars like these.

Headlines for search and ads

Search titles work differently. Google shows one version of a page, so you can't split visitors on a single URL. Test title changes as page-group tests instead. Change the titles on a set of similar pages, such as 20 product pages, and compare their clicks and revenue against 20 similar pages you left alone. Outtest runs SEO tests this way for the same reason.

Ad headlines are the easiest to test, because ad platforms split traffic for you. Judge them on revenue from the people who clicked, not on click-through rate. The Meta ads testing guide covers the details.

Tests to try, and what to measure

  1. Outcome headline vs feature headline. Measure revenue per visitor.
  2. Broad audience vs your best-paying segment named in the headline. Measure revenue per visitor and the share of buyers from that segment.
  3. Headline with a specific number vs without. Measure revenue per visitor.
  4. Pain-led vs outcome-led. Measure revenue per visitor and refunds, since fear-based promises can attract poor-fit buyers.
  5. Headline plus a matching subheadline vs the headline alone. Measure revenue per visitor.
  6. Headline that names the price ("Invoicing from $12 a month") vs one that doesn't. Measure revenue per visitor and plan mix.
  7. Question headline vs statement. Measure revenue per visitor.
  8. For ads, the landing page headline copied from the ad vs your standard headline. Measure revenue per visitor on paid traffic.

Headlines are one step in a wider order of tests. Landing page A/B testing covers where they fit, and revenue per visitor explains why it is the number to judge them on.

Questions people ask

How many headlines should I test at once?+

Two to four. Each extra version splits your traffic further, so the test takes longer to reach an answer. With a few thousand visitors a week, two versions is usually the most you can afford. Make each one a genuinely different promise so the test can find a real difference.

Should I test the headline and subheadline together?+

Usually, yes. Visitors read them as one message, and a new promise in the headline often needs a new subheadline to support it. Test them as a package, and split them apart later only if you need to know which part did the work.

Can I A/B test a page title for Google?+

Not on a single page, because search engines only see one version of a URL at a time. Test titles as a page-group test instead: change the titles on a set of similar pages and compare their search clicks and revenue against similar pages you left alone.

What metric should a headline test use?+

Revenue per visitor wherever the page leads to a purchase. A headline that promises more can win signups from people who won't pay, and a narrower headline can win fewer signups and more revenue. Clicks and signups are useful context, not the decision.

Read next

Let Outtest run your split tests

AI agents read your analytics and payments, find where you lose the most money, build the fix and test it. Every test is judged on revenue, not clicks. Plans from $29 a month.