How many visitors do you need for an A/B test?
The A/B test sample size formula, a worked example and a lookup table. At a 3% conversion rate, detecting a 20% lift needs about 13,900 visitors per version.
Updated 28 September 2026 · 5 min read
It depends on your conversion rate and the smallest lift you want to detect. At a 3% conversion rate, 95% confidence and 80% power, you need about 13,900 visitors per version to detect a 20% relative lift and about 53,200 per version for a 10% lift. Halving the lift you want to detect roughly quadruples the sample, so low-traffic sites should test bold changes.
The number of visitors an A/B test needs depends on four inputs: your current conversion rate, the smallest lift you want to be able to detect, the confidence level (usually 95%) and the power (usually 80%). With those fixed, the answer is a formula, not a judgment call.
Two numbers to anchor on. At a 3% conversion rate, detecting a 20% relative lift (3.0% to 3.6%) takes about 13,900 visitors per version, and detecting a 10% lift (3.0% to 3.3%) takes about 53,200 per version. The /tools/ab-test-sample-size-calculator gives the exact number for your inputs.
The four inputs
- Baseline conversion rate (p₁): what the current version converts at today. Take it from your analytics for the page and audience the test will cover.
- Minimum detectable effect: the smallest relative lift you care about. If a 5% lift wouldn't change what you do, don't size for 5%. /guides/minimum-detectable-effect covers how to choose it.
- Confidence level, usually 95%: how strict you are about false wins. At 95%, a test with no real difference produces a significant result 5% of the time.
- Power, usually 80%: the chance you detect a real lift of the size you chose. At 80%, one in five real lifts of that size will be missed.
The formula
For two conversion rates, the standard formula for visitors per version is:
n = (z₁₋α/₂ × √(2p̄(1 - p̄)) + z₁₋β × √(p₁(1 - p₁) + p₂(1 - p₂)))² / (p₂ - p₁)²
where:
- p₁ is the baseline rate and p₂ = p₁ × (1 + lift) is the rate you want to detect
- p̄ = (p₁ + p₂) / 2
- z₁₋α/₂ = 1.96 for 95% confidence (two-sided)
- z₁₋β = 0.8416 for 80% power
There's also a shortcut. At 95% confidence and 80% power, the formula simplifies to roughly:
n ≈ 16 × p(1 - p) / δ²
where δ is the absolute difference you want to detect (p₂ - p₁). This rule of thumb appears in Ronny Kohavi's Controlled experiments on the web and in Evan Miller's How Not To Run an A/B Test. The 16 comes from (1.96 + 0.84)² × 2 ≈ 15.7, rounded up.
A worked example
A SaaS signup page converts 3% of visitors. The team wants to detect a 10% relative lift, from 3.0% to 3.3%.
With the shortcut:
- p(1 - p) = 0.03 × 0.97 = 0.0291
- δ = 0.033 - 0.030 = 0.003, so δ² = 0.000009
- n ≈ 16 × 0.0291 / 0.000009 = 51,733 per version
With the full formula:
- p̄ = 0.0315, so 1.96 × √(2 × 0.0315 × 0.9685) = 1.96 × 0.2470 = 0.4841
- 0.8416 × √(0.03 × 0.97 + 0.033 × 0.967) = 0.8416 × 0.2470 = 0.2079
- (0.4841 + 0.2079)² = 0.4789
- n = 0.4789 / 0.000009 ≈ 53,211 per version
So about 53,200 visitors per version, or 106,400 in total. The shortcut came within 3%, which is close enough for planning.
If the page gets 2,000 visitors a day, the test needs 53 days. That's too long (see /guides/how-long-to-run-an-ab-test for why tests should end within about six weeks). The team has three choices: test a bolder change and size for a 20% lift (13,914 per version, 14 days), accept that the test can only detect larger lifts, or test a busier page.
Lookup table
Visitors needed per version at 95% confidence and 80% power:
| Baseline rate | 5% lift | 10% lift | 20% lift | 30% lift |
|---|---|---|---|---|
| 1% | 637,010 | 163,095 | 42,693 | 19,827 |
| 2% | 315,206 | 80,682 | 21,109 | 9,798 |
| 3% | 207,938 | 53,211 | 13,914 | 6,455 |
| 5% | 122,124 | 31,234 | 8,158 | 3,780 |
| 10% | 57,763 | 14,751 | 3,841 | 1,774 |
| 20% | 25,583 | 6,510 | 1,683 | 772 |
Two patterns stand out. Halving the lift roughly quadruples the sample, because the difference is squared in the formula. And higher baseline rates need far fewer visitors, which is why tests near the payment step, where conversion rates are high, often finish faster than homepage tests (/guides/what-to-ab-test-first has an example).
What changes the number
| Change | Effect on sample size (3% baseline, 10% lift) |
|---|---|
| The standard settings | 53,211 per version |
| Detect a 20% lift instead | 13,914 (about a quarter) |
| 90% power instead of 80% | 71,233 (about a third more) |
| 99% confidence instead of 95% | 79,177 (about half as many again) |
| Four versions instead of two | Each group grows, and you need a stricter bar per comparison |
The split matters too. A 50/50 split reaches an answer fastest. Kohavi and colleagues note that running a test at 99/1 takes about 25 times longer than at 50/50 to reach the same power.
Revenue needs more visitors
The formula above is for conversion rates, where each visitor either converts or doesn't. Revenue per visitor is different. Most visitors spend nothing, and a few spend a lot. That spread makes the average noisier, so revenue tests need more visitors to detect the same relative lift.
How much more depends on how uneven your order values are. In Seven Rules of Thumb for Web Site Experimenters, Kohavi and colleagues report that revenue per user at Bing was so skewed that they needed at least 114,000 users per version before the average behaved normally. Capping revenue at $10 per user per week cut that to about 9,700, and the capped metric could detect a change 30% smaller with the same sample.
The practical advice is to judge on revenue per visitor anyway, since that's what pays the bills (/guides/revenue-per-visitor explains why), but budget more traffic, and consider capping extreme orders when the tool you use allows it.
Mistakes that break the calculation
- Mixing up relative and absolute lift. Going from 3.0% to 3.6% is a 20% relative lift and a 0.6 percentage point absolute lift. Putting "20" into a calculator that expects absolute points gives nonsense.
- Reading the per-version number as the total. Double it for a two-version test.
- Using the conversion rate of all site visitors when only some will see the test. If the test runs on the pricing page, use the pricing page's rate and traffic.
- Sizing the test after it ends. Power calculated from the observed result is almost meaningless. Kohavi, Deng and Vermeer's A/B Testing Intuition Busters shows a real test with 82 and 75 visitors whose power before it ran was about 3%, yet it reported a significant 337% lift.
- Stopping when significance appears before the planned sample. That turns a 5% false win rate into something much higher (/guides/peeking-problem-ab-testing).
When you can't reach the number
Flip the question. Instead of asking how many visitors you need, ask what lift you can detect with the visitors you have. At a 3% conversion rate, 14,000 visitors per version (1,000 a day for four weeks) can reliably detect a lift of about 20%. The /tools/minimum-detectable-effect-calculator does this for your numbers, and /guides/ab-testing-low-traffic covers what to test when the answer is uncomfortably large.
Outtest's third bar, the minimum lift a winner must show, defaults to 10% and can be set anywhere from 2 to 30% in Settings. Set it with this table in mind. The lower the bar, the more visitors each test needs before it can clear it.
Questions people ask
What is a good sample size for an A/B test?+
There's no single number. It comes from your baseline conversion rate, the smallest lift worth detecting, and your confidence and power settings. At a 3% conversion rate and the usual 95% confidence and 80% power, a 20% lift needs about 13,900 visitors per version and a 10% lift about 53,200.
Is 1,000 visitors enough for an A/B test?+
Only for very large effects or high conversion rates. With 500 visitors per version at a 3% conversion rate, only a lift of more than 100% (3% to about 6.8%) is reliably detectable. At a 50% conversion rate, such as a checkout step, 1,000 visitors in total can detect a lift of about 18%.
Is the sample size per version or in total?+
Calculators usually give it per version. For a two-version test, double it for the total. For a test with four versions, multiply by four, and use a stricter significance bar to account for the extra comparisons.
What is statistical power in A/B testing?+
Power is the probability that your test detects a real lift of the size you chose. At 80% power, one in five real lifts of that size will be missed. Raising power to 90% needs about a third more visitors.
Do revenue tests need a bigger sample than conversion tests?+
Usually yes. Revenue per visitor varies more than a yes or no conversion because order values differ, and a few big orders can swing the average. Capping extreme values or using a longer test helps.
Read next
Let Outtest run your split tests
AI agents read your analytics and payments, find where you lose the most money, build the fix and test it. Every test is judged on revenue, not clicks. Plans from $29 a month.