outtest
Get started
Basics

Is split testing worth it? How to work out the ROI

Split testing pays when the revenue through tested pages is big enough that a few real lifts beat tool and time costs. Here's the formula and a worked example.

Updated 28 September 2026 · 6 min read

The short answer

Split testing is worth it when the revenue flowing through the pages you test is large enough that a few kept lifts outweigh the cost of tools and time. Work out your break-even lift (monthly testing cost divided by monthly revenue through tested pages), then compare it with realistic results: most tests don't win, and winners usually earn less than their test reported. A business with $60,000 a month through its pricing and landing pages and $649 a month in costs breaks even at a 1.1% lift.

Free tool: Split test ROI calculator. No signup.

Split testing is worth it when the revenue that flows through the pages you test is big enough for a handful of real lifts to pay for the tools and the time. The quickest check is your break-even lift: monthly testing costs divided by monthly revenue through the tested pages. If you spend $649 a month on testing and $60,000 a month comes through your landing and pricing pages, you need about a 1.1% lasting lift to break even. If only $5,000 a month comes through, you need 13%.

Then be realistic about returns. Most tests don't win, and winners usually earn less in practice than their test reported. A good ROI estimate uses a low win rate and a discounted lift, and counts each winner only for the months it's actually live.

The evidence that testing pays

The best-known case comes from Microsoft. In 2012, a Bing employee's idea to change how ad headlines were displayed sat unprioritized for more than six months. When someone finally tested it, it raised revenue by 12%, which "on an annual basis would come to more than $100 million in the United States alone". Kohavi and Thomke also report that Microsoft, Amazon, Booking.com, Facebook and Google each run "more than 10,000 online controlled experiments annually", and that testing helped Bing find changes that together raised revenue per search by 10 to 25% a year (Harvard Business Review, 2017).

For smaller companies, the strongest evidence is a study of 35,262 high-tech startups founded between 2008 and 2013, by Rembrand Koning, Sharique Hasan and Aaron Chatterji. The 4,645 startups that adopted A/B testing tools between 2015 and 2018 had, on average, roughly 10% more weekly page views, a 5% greater likelihood of raising venture capital, and launched 9 to 18% more products. They also reached their end points faster. Adopters were more likely to shut down, and more likely to reach well over 50,000 page views a week. The effects often didn't appear until a startup had been testing for six months (Harvard Business School Working Knowledge, 2021).

That study comes with two cautions. It measured page views, funding and launches, not revenue. And companies chose to adopt testing, so the researchers compared each firm with itself before and after adoption to reduce selection bias, which helps but can't remove it.

The ROI formula

Annual ROI = (value of kept lifts − annual testing cost) ÷ annual testing cost

Each part needs an honest input:

Input What it means Where to get it
Revenue through tested pages Monthly revenue from visitors who pass through the pages you'll test Payment tool, filtered by landing or pricing page visitors
Tests per year How many tests your traffic can finish Your minimum detectable effect and test length
Win rate Share of tests that produce a winner you ship Your history, or 10 to 33% from published figures
Realized lift per winner What a winner earns after launch, not what the test showed A holdout group, or discount test results by about half
Months live How long each winner earns in the period Winners ship through the year, so about half the year on average
Annual cost Tools, people's time, design and development Invoices and timesheets

Win rates matter most. Kohavi and Thomke report that at Google and Bing "only about 10% to 20% of experiments generate positive results", and that at Microsoft about a third are positive, a third neutral and a third negative. Small businesses testing bolder changes may do better, but plan on the low end.

Worked example

A B2B software company sends $60,000 a month through its landing and pricing pages. It plans to run two tests a month. All figures below are example inputs.

Costs:

  • Testing tool: $49 a month (Outtest's Growth plan, which covers 10 tests a month, as an example)
  • Time: 6 hours a month to review ideas, check copy and read results, at $100 an hour = $600
  • Total: $649 a month, or $7,788 a year

Value:

  • 24 tests a year at a 20% win rate gives about 5 winners.
  • Each winner's test showed a +6% lift in revenue per visitor. Discount by half for lucky readings and overlap: +3% realized.
  • One realized 3% lift is worth $60,000 × 0.03 = $1,800 a month.
  • Winners arrive through the year, so each is live about 6 months in year one: 5 × $1,800 × 6 = $54,000.

Year-one ROI = ($54,000 − $7,788) ÷ $7,788 = 5.9, or about 590%.

In year two, if those lifts hold and you ran no more tests, the same five winners would be worth 5 × $1,800 × 12 = $108,000. Treat that as an upper bound, because some effects fade over time.

Break-even lift = $649 ÷ $60,000 = 1.08%, as a lasting lift on monthly revenue. One modest real winner covers the year's costs.

Now run the same costs for a business with $5,000 a month through its tested pages. The break-even lift becomes $649 ÷ $5,000 = 13%. And with that little revenue, traffic is probably low too, so the smallest lift the business can detect might be 30% or more. Testing still works there, but only for bold changes, and costs have to come down. Plug your own numbers into the split test ROI calculator.

Costs people forget

  • Development time for each variation, including QA on mobile.
  • Page speed. Client-side testing scripts can add load time. At Bing, Kohavi and Thomke found that every 100 milliseconds of difference in performance had a 0.6% impact on revenue.
  • Revenue lost to losing versions while a test runs. Stopping clear losers early reduces this.
  • The holdout, if you keep one. In the holdout guide's example, a 5% holdout gave up about $1,650 a month against $31,350 of measured gains.
  • Management attention. A test program that nobody reads results from is all cost.

Value people forget

  • Losses avoided. If a third of changes hurt, as Microsoft found, then testing stops you shipping those. The value is real even though it never shows up as a win.
  • Learning. A test that shows annual pricing doesn't change revenue tells you where not to spend the next month.
  • Speed of decisions. In the HBS study, adopters reached scale or shut down sooner, and Koning describes testing as helping founders move on from bad ideas faster.

How to measure the ROI you actually got

Forecasts are guesses. After six months, measure it:

  1. Keep a small holdout (5 to 10% of visitors) on the original experience after each win.
  2. Compare revenue per visitor between the holdout and everyone else, using data from your payment tool, not signups. See revenue per visitor.
  3. Multiply the gap by the visitors on the winning experience to get monthly value.
  4. Subtract the actual costs for the period.

Outtest keeps a 5% holdout after every win for this reason, so it can show what winners actually earned rather than what the tests predicted.

When split testing isn't worth it

  • Your break-even lift is higher than any lift your traffic can detect in six to eight weeks.
  • You're about to change the product, price or positioning, which would make results stale.
  • Nobody on the team will act on results.

In those cases, put the time into getting more traffic, talking to customers, or fixing obvious problems, and come back to testing when volume grows. For where to start when you do, see what to A/B test first and A/B testing with low traffic.

Questions people ask

What ROI do A/B tests typically have?+

There's no reliable industry average, because it depends on revenue, traffic and costs. The best-known examples are large: one Bing ad headline test raised revenue 12%, worth over $100 million a year in the US. For a small business, work it out from your own numbers with the break-even lift and a realistic win rate.

How many A/B tests win?+

Fewer than most people expect. Kohavi and Thomke report that only about 10 to 20% of experiments at Google and Bing generate positive results, and at Microsoft roughly a third are positive, a third neutral and a third negative. Plan your ROI on a low win rate.

Is split testing worth it for a small startup?+

Often not for small tweaks. With low traffic you can only detect big lifts, and with low revenue even real lifts are worth little in dollars. Early on, bold tests of offers, pricing and positioning pay off better, and a study of 35,262 startups found A/B testing adopters launched more products and grew page views faster.

How do I prove testing paid off?+

Keep a small holdout group on the original experience after you ship winners, then compare revenue per visitor between the holdout and everyone else. That measures what the winners actually earned, which is usually less than the sum of the test results.

Read next

Let Outtest run your split tests

AI agents read your analytics and payments, find where you lose the most money, build the fix and test it. Every test is judged on revenue, not clicks. Plans from $29 a month.