Why stopping an A/B test early misleads you (the peeking problem)
Checking an A/B test daily and stopping at the first significant result can turn a 5% false win rate into 25% or more. A simulation, sources and fixes.
Updated 28 September 2026 · 6 min read
If you check an A/B test repeatedly and stop the first time it looks significant, the 5% false win rate you think you have becomes several times larger. In a simulation of 20,000 tests where both versions were identical, checking a 0.05 p-value every day for four weeks produced a 'significant' result at some point in 27.8% of tests. Fix it by fixing the sample size and reading the result once, or by using a sequential method built for repeated looks.
Peeking means checking an A/B test while it runs and stopping the moment the result looks significant. It's the most common way good teams ship changes that do nothing. A 95% significance threshold promises that a test with no real difference will look like a win only 5% of the time, but that promise assumes you look once, at a sample size you fixed in advance. Look every day and stop at the first significant reading, and the false win rate climbs to 20 to 30% over a typical test.
The fix is simple to state. Decide the sample size before launch and read the result once, or use a method designed for repeated looks. The rest of this guide shows why peeking breaks the maths, a simulation you can reproduce, what published sources found, and the options.
Why checking often breaks the maths
When two versions perform the same, the gap between them doesn't sit still at zero. It drifts up and down as visitors arrive, like a random walk. Each time you check, there's a chance the drift happens to be large enough to cross the significance line. One check gives chance one opportunity. Twenty checks give it twenty.
Evan Miller explains it with a small example in How Not To Run an A/B Test (2010). Suppose you check after 200 and after 500 visitors. If you only count the result at 500, you get a false positive 5% of the time. But if you stop at 200 whenever that early check is significant, every test that was significant at 200 and would have drifted back by 500 becomes a false win. You have added false positives and removed none.
A p-value is only valid for the sample size you planned. Stopping when it looks good is choosing the sample size based on the answer.
A simulation you can reproduce
We simulated 20,000 A/A tests, where both versions are identical, to measure exactly how often peeking "finds" a winner that isn't there. The setup:
- Both versions convert at 3%.
- Each day, each version gets 1,000 visitors. Draw each version's daily purchases at random from a binomial distribution (1,000 tries, 3% chance each).
- Keep running totals for 28 days.
- At each check, run a two-sided two-proportion z-test on the totals so far and count the test as a false win if p is below 0.05.
- Repeat 20,000 times and count the share of tests that were ever significant at any check.
Any spreadsheet or a few lines of Python will reproduce it. The results:
| How you check | Tests that looked significant at least once |
|---|---|
| Once, at the end of a 28-day test | 5.0% |
| Once a week for 4 weeks (days 7, 14, 21, 28) | 12.4% |
| Every day of a 7-day test | 17.0% |
| Every day of a 14-day test | 22.3% |
| Every day of a 21-day test | 25.5% |
| Every day of a 28-day test | 27.8% |
Checking once gives the promised 5%. Checking daily for four weeks gives 27.8%, more than five times as many false wins. About half of those false results favor B (13.9% of tests) and half favor A, so if you only act on "B wins", you'd ship a change that does nothing in about one test in seven.
Weekly checks are better but still more than double the error. Every extra look costs something.
What published sources found
The simulation matches what others have reported with different setups:
- Evan Miller's worst case: a test at a 50% conversion rate, checked after every visitor and stopped at the first significant result or at 150 visitors, produced a false positive 26.1% of the time. He also showed that if you peek 10 times, what you think is 1% significance is actually 5% (How Not To Run an A/B Test).
- Optimizely simulated no-difference tests of 5,000 visitors. More than 57% falsely declared a winner or loser at least once when checked after every visitor, 26% when checked every 500 visitors and 20% when checked every 1,000 (Optimizely, 2015).
- David Robinson simulated 20-day tests with about 10,000 impressions a day and no real effect. 5.34% were significant on day 20, but 22.68% dipped below p = 0.05 at some point along the way (Is Bayesian A/B Testing Immune to Peeking? Not Exactly, 2015).
- Kohavi, Deng and Vermeer note in A/B Testing Intuition Busters that an early commercial A/B platform showed near-real-time results and recommended stopping when results turned significant, which inflated false positive rates.
The problem was described for medical trials long before web testing. Evan Miller's article points to Armitage, McPherson and Rowe's 1969 paper "Significance Tests on Accumulating Data".
Bayesian methods don't escape it
Switching to a Bayesian "chance to beat" doesn't remove the problem if you use it as a stopping rule. Running the same 20,000 A/A tests and stopping when B's chance to beat first reached 95% (flat prior, normal approximation):
- Checked once, on day 28: 5.1% declared B the winner
- Checked every day: 23.6%
Robinson found the same with a Bayesian expected-loss rule. Checking daily raised the rate of switching to a worse version from 2.5% to 11.8%. /guides/bayesian-vs-frequentist-ab-testing explains when a Bayesian posterior does stay valid under early stopping (short version: when the prior is a real, well-founded belief).
Extending a test is peeking too
The mirror-image mistake is letting a test run past its planned end because it's "almost significant". Adding days until the result crosses the line gives chance extra attempts in the same way. If the test reaches its planned sample without clearing the bar, it's a draw.
When stopping early is fine
Stopping to protect the business is fine. If a version breaks checkout, throws errors or is clearly losing revenue, stop it. The peeking problem is about declaring winners early, not about pulling a broken version. A "loser" call made early is still at risk of being noise, but the cost of wrongly stopping a bad-looking version is usually small. You keep the original.
How to fix it
| Approach | How it works | Trade-off |
|---|---|---|
| Fixed sample, one look | Calculate the sample size, run whole weeks, analyze once at the end | Simple and valid; no early stopping for big wins |
| A few planned checkpoints | Decide in advance to look, say, 5 times, and use a stricter bar at each look | Evan Miller's table shows you need about 1.4% per look to keep 5% overall with 5 looks |
| Sequential testing | Always valid p-values or a simple sequential rule that allows a look after every visitor | Needs a larger maximum sample; you must use the method's own thresholds |
| Bayesian with an informed prior | Priors built from past tests keep early results from running away | Only as good as the prior |
| Guard bars | A minimum run time, a minimum lift and a maximum duration | Reduce the damage but don't restore an exact 5% |
For sequential methods, see Optimizely's description of its sequential engine above, the always valid inference paper by Johari, Pekelis and Walsh, and Evan Miller's Simple Sequential A/B Testing, which you can run by hand. That rule: choose a number of conversions N before the test. Stop and declare B the winner if B's conversions minus A's reach 2√N. Stop with no winner once the two groups' conversions add up to N.
A peeking-safe routine
For most teams, this routine avoids the problem without new software:
- Before launch, calculate the sample size and round the duration up to whole weeks with the /tools/ab-test-duration-calculator.
- Write the end date where the team can see it.
- During the test, check only health: the traffic split (/guides/sample-ratio-mismatch), errors, and whether a version is breaking something.
- On the end date, read the result once and act on it. A result short of the bar is a draw.
- After shipping a winner, keep a small holdout on the original to confirm the lift was real.
Outtest's Referee won't call a winner before a test has run at least 7 days, and ends any test that hasn't cleared its bars by day 42 as a draw, so a test can't run indefinitely waiting for a lucky day. After a win, 5% of visitors keep seeing the original, so a lucky call shows up as a winner that earns no more than the version it replaced.
/guides/how-long-to-run-an-ab-test covers how to set the end date in the first place.
Questions people ask
What is peeking in A/B testing?+
Peeking means checking a running test's significance repeatedly and stopping as soon as it crosses the threshold. Each check gives chance another opportunity to produce a significant-looking gap, so the real false win rate ends up far above the 5% a 95% threshold promises.
Is it OK to look at A/B test results while the test is running?+
Looking is fine. Acting on it is the problem. Watch for bugs, a broken traffic split or a version that is clearly losing money, and stop for those. Don't declare a winner before the planned sample size unless you're using a sequential method designed for early stopping.
How much does peeking inflate false positives?+
It depends on how often you look. In our simulation of identical versions checked for four weeks, one check at the end gave a 5.0% false win rate, weekly checks gave 12.4% and daily checks gave 27.8%. Evan Miller showed that checking after every visitor can reach 26.1% in a short test, and Optimizely reported more than 57% in simulated 5,000-visitor tests checked after every visitor.
Does Bayesian A/B testing solve the peeking problem?+
Not on its own. Stopping when a Bayesian chance to beat first crosses 95% also inflates false wins: in our simulation it went from 5.1% with one check to 23.6% with daily checks. A Bayesian posterior stays calibrated under early stopping only if the prior is a genuine belief that matches reality.
Read next
Let Outtest run your split tests
AI agents read your analytics and payments, find where you lose the most money, build the fix and test it. Every test is judged on revenue, not clicks. Plans from $29 a month.