Statistical significance in A/B testing, explained
What statistical significance and p-values mean in an A/B test, what 95% does and doesn't tell you, and how often significant winners turn out to be false.
Updated 28 September 2026 · 6 min read
An A/B test result is statistically significant when its p-value falls below a bar set in advance, usually 0.05. The p-value is the chance of seeing a gap at least as big as yours if the two versions truly performed the same. It is not the chance that B is better, and a significant result is not 95% certain: when only about 10% of ideas work, roughly 1 in 5 significant winners is a false positive even in a well-run test.
An A/B test result is statistically significant when its p-value is below a threshold you chose before the test, almost always 0.05. The p-value answers one narrow question: if the two versions truly performed the same, how often would chance alone produce a gap at least as big as the one you saw? A p-value of 0.03 means "about 3% of the time".
That's all it means. It is not the probability that version B is better, and "95% significant" does not mean you're 95% likely to be right. How often a significant winner is real depends on how good your ideas are and how well the test was sized, and in well-run programs a meaningful share of significant winners are still false.
The p-value in plain words
Imagine every visitor flips a slightly weighted coin that lands "buy" 3% of the time, and that your change did nothing, so both versions use the same coin. Split 20,000 flips into two groups of 10,000 and count the buys. The two counts won't match exactly. Sometimes one group gets 290 and the other 315 just by luck.
The p-value asks: in that no-difference world, how often would the gap between the groups be as large as the one in your real test, or larger? If it would happen often, your result is unremarkable. If it would almost never happen, you have evidence that the versions differ.
Kohavi, Deng and Vermeer give the formal version in A/B Testing Intuition Busters. The p-value is the probability of a result equal to or more extreme than what was observed, assuming every modeling assumption, including "no difference", is true. They stress that the condition "assuming no difference" is the part most often forgotten.
A worked example
A pricing page test, with example numbers:
| Version | Visitors | Purchases | Conversion rate |
|---|---|---|---|
| A | 10,000 | 300 | 3.00% |
| B | 10,000 | 345 | 3.45% |
B shows a 15% lift. Is it significant? A two-proportion z-test:
- Pooled rate: (300 + 345) / 20,000 = 3.225%
- Standard error: √(0.03225 × 0.96775 × (1/10,000 + 1/10,000)) = 0.002498
- z = (0.0345 - 0.0300) / 0.002498 = 1.80
- Two-sided p-value for z = 1.80: 0.072
At the 0.05 bar, this is not significant. The 95% confidence interval for the lift runs from about -1% to +31%. B might be much better, or a little worse.
Now suppose the test had been sized properly and ran to 20,000 visitors per version with the same rates (600 against 690 purchases). The standard error shrinks to 0.001767, z becomes 2.55, and the p-value drops to 0.011. The interval narrows to about +3.5% to +26.5%. Same lift, more data, a clear answer.
You can check any result like this in the /tools/ab-test-significance-calculator.
What 95% does mean
Using a 0.05 threshold is a promise about the long run. If you run many tests where the new version truly makes no difference, and you analyze each once at its planned end, about 1 in 20 will come out "significant" by chance. The threshold caps how often you ship a change that did nothing, among tests of changes that did nothing.
That promise only holds when the test follows the rules: sample size fixed in advance, one planned analysis, one pre-chosen metric. Check daily and stop at the first significant reading, and the false win rate climbs well past 5% (/guides/peeking-problem-ab-testing shows by how much).
What 95% doesn't mean
The American Statistical Association's 2016 statement on p-values lists six principles. Two matter most for A/B testers: p-values do not measure the probability that the studied hypothesis is true, and a p-value does not measure the size of an effect or the importance of a result.
So a p-value of 0.03 does not mean:
- A 97% chance that B is better than A.
- A 3% chance that the result is a fluke.
- That the lift you measured is the lift you'll get.
- That the lift is large enough to matter.
The first two mistakes are common enough that Intuition Busters calls them out in vendor documentation and textbooks. The authors note that "confidence", shown by many tools as (1 - p-value) × 100%, is often misread as the probability that the result is a true positive.
How often significant winners are false
The chance that a significant result is a false positive depends on how many of your ideas work in the first place. Most don't. Kohavi and Thomke report in HBR that only about 10 to 20% of experiments at Google and Bing generate positive results.
Intuition Busters works out the false positive risk from published success rates, for tests at 80% power:
| Organization | Success rate | Share of significant wins that are false |
|---|---|---|
| Microsoft | 33% | 5.9% |
| Bing | 15% | 15.0% |
| Booking.com, Google Ads, Netflix | 10% | 22.0% |
| Airbnb search | 8% | 26.4% |
At 20% power, the Booking.com row rises to 52.9%. Underpowered tests are where most false winners come from.
The maths is Bayes' rule. Take 1,000 ideas where 10% truly work. With 80% power you catch 80 of the 100 real winners. Of the 900 duds, a one-sided 2.5% bar lets through about 22. So 22 of your 102 "wins", about 22%, are false. The authors recommend a stricter bar for surprising results and replicating them before trusting them.
Confidence intervals say more than p-values
A 95% confidence interval gives the range of lifts that fit your data. It tells you three things a p-value can't:
- The direction: an interval entirely above zero means B is very likely better.
- The size: "+3.5% to +26.5%" is a very different business case from "+15% to +16%".
- The uncertainty: a wide interval means you learned less than the headline lift suggests.
When the interval crosses zero, the test hasn't ruled out that B is no better, or worse.
Statistical significance isn't business significance
With a million visitors per version, a lift of 0.5% can be highly significant. Whether it's worth shipping depends on the cost of building and maintaining the change, and on whether the metric was revenue or something easier to move. Decide the smallest lift worth acting on before the test (/guides/ab-test-sample-size) and treat anything smaller as a draw even if it's significant.
One-sided or two-sided
A two-sided test asks "is B different from A, in either direction?". A one-sided test asks only "is B better?". A one-sided p-value is half the two-sided one, which makes significance easier to reach, so choose before the test, not after. Two-sided is the safer default because it also flags versions that hurt you.
Significance and "chance to beat"
Many tools show a Bayesian "chance B beats A" instead of, or next to, a p-value. With a flat prior and plenty of data, the two line up closely. In the example above, the one-sided p-value of 0.036 corresponds to a chance to beat of about 96%. The Bayesian number is easier to explain, but it has the same weaknesses if you stop early or run underpowered tests. /guides/bayesian-vs-frequentist-ab-testing compares the two.
Outtest calls winners with a Bayesian chance to beat on revenue per visitor, and requires three bars together: at least 7 days, at least a 90% chance of beating the original, and a lift of at least 10%. Each bar can be adjusted in Settings.
Before you trust a significant result
- The sample size was set before launch and reached.
- The test ran whole weeks.
- The traffic split matches what you configured (/guides/sample-ratio-mismatch).
- The decision metric was chosen before launch.
- The lift is plausible. If it looks too good, check the tracking first.
- The confidence interval clears the smallest lift worth shipping.
Questions people ask
What does 95% statistical significance mean in an A/B test?+
It means the test used a 0.05 threshold: if the two versions truly performed the same, a difference as large as the one observed would appear less than 5% of the time. It controls how often you declare a winner when nothing changed. It does not mean there's a 95% chance version B is better.
What is a good p-value for an A/B test?+
Below 0.05 is the common bar, set before the test starts. For surprising results or expensive decisions, a stricter bar like 0.01 is safer. A p-value only counts if the test ran to its planned sample size without being stopped early.
Can a test be significant but not important?+
Yes. With enough traffic, a 0.5% lift can be highly significant and still not be worth the cost of shipping and maintaining the change. Look at the size of the lift and its confidence interval, not just the p-value.
Why did my test show a 15% lift but not reach significance?+
Because the sample was too small to rule out chance. With 10,000 visitors per version at a 3% conversion rate, a 15% lift gives a p-value of about 0.07 and a confidence interval that includes zero. The same lift on twice the traffic would be significant.
Read next
Let Outtest run your split tests
AI agents read your analytics and payments, find where you lose the most money, build the fix and test it. Every test is judged on revenue, not clicks. Plans from $29 a month.