outtest
Get started
Basics

Common A/B testing mistakes (and how to avoid them)

The A/B testing mistakes that produce fake winners: stopping early, tiny samples, the wrong metric and broken splits. What each costs and how to fix it.

Updated 28 September 2026 · 7 min read

The short answer

Most bad A/B test results come from a handful of mistakes: stopping the moment a result looks significant, running tests too small to detect anything, judging on clicks or signups instead of revenue, and missing a broken traffic split. Stopping at the first significant result can push a 5% false win rate above 25%. Fix them by sizing the test first, running whole weeks, judging on revenue per visitor and checking the split.

Free tool: A/B test significance calculator. No signup.

The mistakes that do the most damage in A/B testing all end the same way, with a winner that isn't real. The big ones are stopping a test the moment it looks significant, running tests too small to detect the lift you're hoping for, judging the test on clicks or signups instead of money, and trusting a test whose traffic split is broken.

Each has a simple fix, and none needs a statistician. Below are thirteen mistakes grouped by when they happen, with what each one costs and what to do instead.

Before the test

1. Running a test without calculating the sample size

Without a sample size you don't know whether the test can detect anything. The result is noise that looks like a finding.

Ronny Kohavi, Alex Deng and Lukas Vermeer open A/B Testing Intuition Busters with a real, publicly shared test titled "Which design radically increased conversions 337%?". It had 82 visitors in the control and 75 in the treatment, with 3 and 12 conversions (3.7% against 16.0%). The p-value was 0.009, which looks convincing. The authors estimate the test's power beforehand was about 3%. Even assuming a third of ideas genuinely work, a significant result from a test that weak has a 63% chance of being a false positive.

Fix it by working out the sample size before launch with the /tools/ab-test-sample-size-calculator. /guides/ab-test-sample-size shows the formula.

2. Chasing small lifts on small traffic

Halving the lift you want to detect roughly quadruples the visitors you need. At a 3% conversion rate, detecting a 20% lift takes 13,914 visitors per version, a 10% lift takes 53,211 and a 5% lift takes 207,938.

If your page gets 1,000 visitors a day at a 3% conversion rate, a four-week test can reliably detect a lift of about 20%, not 5%. Test changes that could plausibly move the number that much: a different offer, a different price, a shorter checkout, a headline that promises something new. /guides/ab-testing-low-traffic has more options.

3. Judging the test on the wrong metric

Clicks, add-to-carts and signups are easy to move without moving revenue. A worked example with made-up numbers: a SaaS pricing page test where version B makes the cheapest plan more prominent.

Version Visitors Signups Revenue Revenue per visitor
A 10,000 400 (4.0%) $12,000 $1.20
B 10,000 480 (4.8%) $10,800 $1.08

B wins signups by 20% and loses 10% of revenue per visitor, because more people chose the cheap plan. If you judged on signups, you'd ship a pay cut.

Fix it by deciding the metric before launch and making it revenue per visitor wherever money changes hands. Outtest judges every website, pricing and ad test on revenue per visitor read from the payment tool and shows signups for context only. /guides/revenue-per-visitor covers the maths, including why revenue needs more visitors than a conversion rate.

4. Changing five things and calling it one test

A new layout, new copy, new photos and a new offer in one version is a redesign test. If it wins, you don't know which part did the work. If it loses, one strong change may have been dragged down by a weak one.

That's fine when you want a fast answer on a package. It's a mistake when you'll later claim "the new headline added 12%". Test one idea per test when you need to know what worked.

During the test

5. Stopping as soon as it looks significant

Significance wanders as data arrives. A test with no real difference will cross the 95% line at some point surprisingly often, and if you stop there, you ship a false winner.

Evan Miller showed that checking after every visitor and stopping at the first significant result pushes the real false positive rate to 26.1%, more than five times the 5% you think you have (How Not To Run an A/B Test). Optimizely's own simulations of 5,000-visitor tests with no real difference found that more than 57% declared a winner or loser at least once when checked after every visitor (Optimizely, 2015).

Fix it by fixing the sample size and end date in advance, or by using a sequential method designed for repeated looks. /guides/peeking-problem-ab-testing shows a simulation and the options.

6. Running less than a full week

Weekday and weekend visitors often behave differently. A test that runs Monday to Thursday measures a different audience from the one you'll ship to. Kohavi and colleagues recommend running for at least a week or two and extending in whole weeks so day-of-week effects even out (Controlled experiments on the web).

Fix it by rounding every test up to whole weeks. /guides/how-long-to-run-an-ab-test covers minimum and maximum run times.

7. Ignoring a broken traffic split

If you set 50/50 and one version has noticeably fewer visitors, something is removing visitors from that version: a redirect that fails, a bot filter, a script that loads late. Those missing visitors aren't random, so the comparison is biased. This is called a sample ratio mismatch.

It's common. About 6% of experiments at Microsoft had one (Fabijan et al., KDD 2019). And small-looking gaps count: a 50.2/49.8 split across 1.6 million users has less than a one in 500,000 chance of happening by luck, as Kohavi and Thomke point out in HBR.

Fix it by running a chi-square check on visitor counts early and again at the end. The /tools/srm-checker does it, and /guides/sample-ratio-mismatch explains the causes.

8. Changing the test while it runs

Editing the new version, changing the traffic split or adding a new audience mid-test mixes two different experiments into one set of numbers. Changing allocation over time can even produce Simpson's paradox, where the combined result points the opposite way from every individual period. Kohavi's team keeps the proportions assigned to each version stable for the whole experiment for this reason (Seven Rules of Thumb).

Fix it by restarting the test if you need to change it.

9. Mistaking a novelty effect for a win

Returning visitors sometimes click on something just because it's new. The effect fades within days or weeks. In the Rules of Thumb paper, one MSN change drew 20% of all feedback messages on launch day, most of them negative, falling to 4% in week two and 2% in weeks three and four. The authors recommend running experiments for two weeks to look for such effects, though they also note novelty effects are uncommon in practice.

Fix it by checking whether the lift shrinks week by week, or by comparing new visitors (who have no habit to break) with returning ones.

After the test

10. Reading "95% significant" as "95% sure B is better"

A p-value below 0.05 does not mean a 5% chance the result is wrong. How often a significant result is false depends on how many of your ideas actually work. Kohavi, Deng and Vermeer estimate this false positive risk. With a 10% success rate, which they cite for Booking.com, Google Ads and Netflix, about 22% of significant results in tests at 80% power are false. At 20% power it's more than half.

Fix it by treating 95% as the minimum, not proof, and replicating surprising results. /guides/statistical-significance-explained goes through it.

11. Hunting through segments and metrics

Check 20 segments (mobile, desktop, new, returning, each country) and the chance that at least one crosses p = 0.05 by luck alone is 1 - 0.95²⁰ = 64%. The same happens with metrics and with data cleaning choices. The Intuition Busters paper cites work showing that removing outliers separately within each version, rather than across all the data, can push false positive rates as high as 43%.

Fix it by naming one decision metric before launch, applying any data cleaning the same way to both versions, and treating segment findings as ideas for the next test.

12. Believing results that look too good

Kohavi and Thomke write in HBR that the best data scientists follow Twyman's law, which says any figure that looks interesting or different is usually wrong. A 60% lift from a button color is far more likely to be a tracking bug, a bot or a broken split than a discovery. Kohavi, Deng and Vermeer write that across tens of thousands of tests at Airbnb, Booking, Amazon and Microsoft, they had never seen a change improve conversions by anywhere near 300%.

Fix it by checking the plumbing first (tracking, split, bots) and rerunning surprising winners before shipping them.

13. Shipping and never checking

A test measures a lift during a few weeks on a sample. Some winners are lucky draws, and lucky draws shrink after launch. If you never check, you'll overestimate what your testing program earns.

Fix it by keeping a small holdout, a few percent of visitors who keep seeing the original, for some weeks after launch. Outtest keeps 5% of visitors on the original after every win so the real earnings of each winner are visible. /guides/holdout-groups explains how to size one.

The short version

Mistake Fix
No sample size Calculate it before launch
Tiny lifts on small traffic Test bolder changes
Clicks or signups as the metric Judge on revenue per visitor
Many changes in one version One idea per test when you need to know why
Stopping at first significance Fixed end date or a sequential method
Partial weeks Round up to whole weeks
Broken split Chi-square check early and at the end
Mid-test edits Restart instead
Novelty effects Compare week by week, new against returning
95% read as certainty Higher bar for surprises, replicate
Segment hunting One pre-set metric, segments become new tests
Too-good results Check tracking, rerun
No follow-up Keep a holdout

Questions people ask

What is the most common A/B testing mistake?+

Stopping a test as soon as it looks significant. Significance swings up and down as data arrives, so checking often and stopping at the first good reading produces far more false winners than the 5% a 95% bar promises. Decide the sample size and end date before launch, or use a sequential method built for continuous monitoring.

Why did my A/B test winner not improve revenue after launch?+

Usually one of three reasons: the win was luck (common in small or stopped-early tests), the test measured clicks or signups rather than revenue, or a novelty effect faded. Keeping a small holdout on the original after launch shows whether the winner is really earning more.

How do I know if my A/B test result is trustworthy?+

Check that the sample size was set in advance and reached, the test ran whole weeks, the traffic split matches what you configured, the decision metric was chosen before launch, and the result isn't implausibly large. A test that fails any of these checks should be rerun, not shipped.

Is it a mistake to test more than one change at once?+

Not always, but you lose the ability to say which change did the work. If the package wins, ship it and test the parts later if you need to know. If it loses, a strong idea may have been dragged down by a weak one.

Read next

Let Outtest run your split tests

AI agents read your analytics and payments, find where you lose the most money, build the fix and test it. Every test is judged on revenue, not clicks. Plans from $29 a month.