Bayesian vs frequentist A/B testing
Frequentist tests give a p-value; Bayesian tests give the chance B beats A. How the two differ, the same data analyzed both ways, and why neither fixes peeking.
Updated 28 September 2026 · 6 min read
Frequentist A/B testing asks how surprising your data would be if there were no difference (the p-value) and controls false wins over many tests with a fixed sample size. Bayesian A/B testing combines the data with a prior belief to give the probability that B beats A and the expected cost of choosing wrong. With a flat prior and decent traffic they usually agree, and neither method makes it safe to stop the moment a result looks good.
Frequentist and Bayesian A/B testing ask different questions of the same data. A frequentist test asks: if the two versions truly performed the same, how surprising is this result? The answer is a p-value, and the method promises that if you fix the sample size in advance, you'll declare a false winner only 5% of the time when nothing changed. A Bayesian test asks: given this data and what I believed beforehand, how likely is it that B is better, and how much could I lose by choosing it? The answer is a "chance to beat" and an expected loss.
In practice, with a neutral prior and a properly sized test, the two usually point the same way. The bigger risks are shared: stopping early, running underpowered tests and judging on the wrong metric break both methods.
The difference at a glance
| Frequentist | Bayesian | |
|---|---|---|
| Main output | p-value, confidence interval | Chance B beats A, credible interval, expected loss |
| Question answered | How surprising is this data if there's no difference? | How likely is B to be better, given the data and prior? |
| Uses a prior belief | No | Yes, flat or informed by past tests |
| Guarantee | Caps false wins at your threshold, if the sample is fixed in advance | Probabilities are calibrated if the prior matches reality |
| Easy to explain | Often misread ("97% sure") | Reads naturally ("96% chance B is better") |
| Early stopping | Breaks the guarantee unless you use a sequential design | Still inflates wrong calls, and depends on the prior |
The same data, analyzed both ways
Here's a pricing page test with example numbers:
| Version | Visitors | Purchases | Conversion rate |
|---|---|---|---|
| A | 10,000 | 300 | 3.00% |
| B | 10,000 | 345 | 3.45% |
Observed lift: 15%.
The frequentist analysis, a two-proportion z-test:
- Pooled rate 3.225%, standard error 0.002498
- z = (0.0345 - 0.0300) / 0.002498 = 1.80
- Two-sided p-value = 0.072, one-sided p-value = 0.036
- 95% confidence interval for the lift: about -1% to +31%
Verdict at the usual 0.05 two-sided bar: not significant. Keep testing to the planned sample, or call it inconclusive.
The Bayesian analysis with a flat prior, Beta(1, 1), meaning every conversion rate starts equally plausible:
- A's posterior: Beta(1 + 300, 1 + 9,700) = Beta(301, 9,701)
- B's posterior: Beta(1 + 345, 1 + 9,655) = Beta(346, 9,656)
- Chance B beats A (by simulation, or with Evan Miller's closed-form formula): 96.4%
- 95% credible interval for the lift: about -1% to +34%
- Expected loss if you choose B: 0.0036 percentage points of conversion. Expected loss if you keep A: 0.45 points.
Verdict at a 95% chance-to-beat bar: B wins. At a 99% bar: not yet.
Notice the one-sided p-value (0.036) and the chance to beat (96.4%) are close to mirror images. With a flat prior and reasonable sample sizes this is typical, and it's why the two schools disagree less in practice than the debate suggests. The real difference in the verdicts comes from the bar each team sets (two-sided 0.05 is stricter than a 95% chance to beat), not from the philosophy.
The prior changes the answer
A flat prior says "any conversion rate is equally likely". But you know more than that. Most changes move conversion by a few percent at most: Kohavi and Thomke report in HBR that only about 10 to 20% of experiments at Google and Bing generate positive results.
A skeptical prior encodes that. Using a normal approximation, suppose the prior says the true difference is centered on zero with a standard deviation of 0.15 percentage points, so lifts bigger than about 10% either way are rare. Combined with the data above:
- Posterior mean lift falls from 15% to about 4%
- Chance B beats A falls from 96.4% to about 82%
Same data, different conclusion. This is the Bayesian method working as designed. A surprising result from a moderate sample gets pulled toward what usually happens. It's also the main criticism of it. If the prior is wrong, the probabilities are wrong, and a vendor's default prior may not match your business.
Large experimentation teams build priors from their own history. Kohavi, Deng and Vermeer note in A/B Testing Intuition Busters that Microsoft's experimentation platform estimates the chance a treatment effect is real using Bayes' rule with priors from historical experiments.
Neither method makes peeking safe
A popular claim is that Bayesian tests let you check results whenever you like and stop as soon as you're confident. Evan Miller wrote in 2015 that with Bayesian formulas you don't need a pre-set sample size. The claim needs a big qualification.
David Robinson tested it in Is Bayesian A/B Testing Immune to Peeking? Not Exactly (2015). In his simulations of a change that actually lowered click-through rate, a Bayesian expected-loss rule switched to the worse version in 2.5% of tests when checked once at the end. When checked daily and stopped early, it switched in 11.8%, more than four times as often.
We ran a similar simulation for this guide: 20,000 A/A tests (no real difference), a 3% conversion rate, 1,000 visitors per version per day, 28 days. Declaring B the winner when its chance to beat (flat prior, normal approximation) reached 95%:
- Checked once, on day 28: 5.1% of tests declared B a winner
- Checked every day, stopping at the first crossing: 23.6%
That's the same inflation you get from peeking at p-values (/guides/peeking-problem-ab-testing has the frequentist version).
The Bayesian side has a precise answer to this. Alex Deng, Jiannan Lu and Shouyuan Chen at Microsoft proved in Continuous monitoring of A/B tests without pain (2016) that a Bayesian posterior stays valid under optional stopping, as long as the prior is a genuine, pre-defined belief. With a proper prior, a 95% posterior still means 95%. What changes is the frequentist error rate, which the Bayesian method never promised to control. With a flat prior, which treats huge lifts as just as plausible as tiny ones, the posterior is only as trustworthy as that assumption.
The frequentist answer to peeking
Frequentists solved early stopping decades ago with sequential designs, which set the rules for repeated looks in advance and adjust the threshold so the overall false win rate stays at 5%.
- Optimizely moved to a sequential frequentist engine in January 2015. Its write-up reports that in simulated no-difference tests checked after every visitor, the false declaration rate fell from 57% with classic statistics to 3% with the sequential method. The underlying idea, always valid p-values, is set out in Johari, Pekelis and Walsh's paper, which says the method was implemented in a large commercial A/B testing platform.
- Evan Miller's Simple Sequential A/B Testing gives a rule you can run by hand: choose a number of conversions N in advance, then stop and declare B the winner if B's conversions minus A's reach 2√N, or stop with no winner once the total reaches N.
The trade-off is that sequential tests need somewhat more data than a fixed-sample test when you run them to the end, in exchange for the right to stop early when an effect is large.
Which should you use?
| If you need | Lean toward |
|---|---|
| Results a non-statistician will read correctly | Bayesian chance to beat, with a stated prior |
| A hard cap on false wins across many tests | Frequentist with a fixed sample, or a sequential design |
| The freedom to stop early on big effects | Sequential frequentist, or Bayesian with an informed prior and a minimum run time |
| Decisions weighted by money at risk | Bayesian expected loss |
| Low traffic | Neither fixes it. Test bigger changes (/guides/ab-testing-low-traffic) |
Whichever you choose, the same rules protect you: decide the metric and the bar before launch, size the test, run whole weeks, check the traffic split and don't treat the first good-looking day as the answer.
Outtest uses a Bayesian chance to beat on revenue per visitor with a normal approximation, and a winner must also clear two other bars: at least 7 days of data and a lift of at least 10%. Tests that don't clear all three by day 42 end as a draw. The /tools/ab-test-significance-calculator shows both the Bayesian chance to beat and the frequentist p-value for any result, so you can see how closely they track on your own numbers.
Questions people ask
Is Bayesian A/B testing better than frequentist?+
Neither is better in general. Bayesian results are easier to explain ('a 96% chance B is better') and handle expected loss naturally. Frequentist methods give clear long-run error guarantees when the sample size is fixed in advance. Discipline about sample size, run time and metric choice matters more than which school you pick.
Can you peek at a Bayesian A/B test?+
You can look, but stopping the moment the chance to beat crosses a threshold still produces more false winners. David Robinson's simulations found that a Bayesian expected-loss rule picked a worse version 2.5% of the time when checked once at the end, and 11.8% when checked daily and stopped early.
What is 'chance to beat' in A/B testing?+
It's the Bayesian probability that version B's true conversion rate (or revenue per visitor) is higher than A's, given the data and the prior. A 95% chance to beat means that, under the model, B is better in 95% of the plausible worlds consistent with what you observed.
Do Bayesian A/B tests need a smaller sample size?+
Not by themselves. The data carries the same information either way. A Bayesian test can reach a decision sooner if you accept a lower bar or use an informative prior, but those choices trade speed for more wrong calls, just as a looser p-value bar would.
Read next
Let Outtest run your split tests
AI agents read your analytics and payments, find where you lose the most money, build the fix and test it. Every test is judged on revenue, not clicks. Plans from $29 a month.