outtest
Get started
Statistics

Minimum detectable effect (MDE), explained

MDE is the smallest lift your A/B test can reliably detect with its traffic. Learn the formula, a worked example, and how to pick an MDE before you launch.

Updated 28 September 2026 · 5 min read

The short answer

The minimum detectable effect (MDE) is the smallest true lift your test has an 80% chance of detecting at 95% confidence, given your traffic and baseline conversion rate. The quick formula is MDE ≈ 4 × √(p(1 − p) ÷ n), with n visitors per version. With 10,000 visitors a week, a 3% conversion rate and a 4-week test, the MDE is about 0.48 percentage points, a 16% relative lift, so smaller real gains will usually go unnoticed.

Free tool: Minimum detectable effect calculator. No signup.

The minimum detectable effect (MDE) is the smallest true improvement your A/B test is likely to catch. By convention, "likely" means an 80% chance (the test's power) of a significant result at 95% confidence. It depends on your traffic per version, your baseline conversion rate, and the confidence and power you choose. If the real lift is smaller than your MDE, your test will usually end without a winner, even though the change helped.

For a conversion rate, the quick formula is MDE ≈ 4 × √(p(1 − p) ÷ n), where p is the baseline rate and n is visitors per version. A site with 10,000 visitors a week, a 3% conversion rate and a four-week test has an MDE of about 0.48 percentage points, or 16% in relative terms. Knowing that before launch tells you which ideas are worth testing and which will end in noise.

Where the formula comes from

The standard sample size formula for comparing two conversion rates, at 95% confidence (two-sided) and 80% power, is roughly:

n = 16 × p(1 − p) ÷ Δ²

where n is visitors per version and Δ is the absolute difference you want to detect. It's given in Kohavi and colleagues' practical guide to controlled experiments, which notes that you replace 16 with 21 for 90% power. The 16 comes from 2 × (1.96 + 0.84)² = 15.7, where 1.96 and 0.84 are the normal-distribution values for 95% confidence and 80% power.

Solve for Δ and you get the MDE:

Δ = √(16 × p(1 − p) ÷ n) = 4 × √(p(1 − p) ÷ n)

Divide by p to get the relative MDE.

Settings Constant (replaces 16) Multiplier in the MDE formula
95% confidence, 80% power 15.7 (use 16) 4.0
95% confidence, 90% power 21.0 4.6
90% confidence, 80% power 12.4 3.5

Worked example

An online course business gets 10,000 visitors a week to its sales page, and 3% of them buy. It splits traffic 50/50, so each version gets 5,000 visitors a week. Here p(1 − p) = 0.03 × 0.97 = 0.0291.

For a four-week test, n = 20,000 per version:

  • Absolute MDE = 4 × √(0.0291 ÷ 20,000) = 4 × 0.00121 = 0.0048, or 0.48 percentage points.
  • Relative MDE = 0.48 ÷ 3.0 = 16%.

So the test can reliably detect a move from 3.0% to about 3.48%. A real lift to 3.2% (7% relative) would most likely finish as "no significant difference".

Test length Visitors per version Absolute MDE Relative MDE
2 weeks 10,000 0.68 points 22.7%
4 weeks 20,000 0.48 points 16.1%
6 weeks 30,000 0.39 points 13.1%
8 weeks 40,000 0.34 points 11.4%
12 weeks 60,000 0.28 points 9.3%

Doubling the test length doesn't halve the MDE. It shrinks it by a factor of √2, about 29%. To halve the MDE you need four times the traffic. That's why low-traffic sites hit a wall. Past six to eight weeks, extra time buys very little. Try your own numbers in the minimum detectable effect calculator.

Absolute vs relative MDE

It's easy to mix these up. A "5% MDE" could mean 3.0% to 3.15% (relative) or 3.0% to 8.0% (absolute), and those need wildly different traffic. Always state both the baseline and which kind you mean. Relative is more common in tools and reports, and it's what this guide uses unless it says "points".

The same absolute lift is also harder to detect at a higher baseline, because p(1 − p) gets bigger up to 50%. But the relative MDE usually falls as the baseline rises. Testing an email signup with a 20% rate is easier than testing purchases at 2%, which is one reason low-traffic teams often test further up the funnel.

MDE for revenue per visitor

Revenue is noisier than conversion, because most visitors spend nothing and a few spend a lot. For a continuous metric, replace p(1 − p) with the variance σ², so the MDE is 4 × σ ÷ √n in dollars.

Kohavi and colleagues give an example: a store where 5% of visitors buy and spend about $75 has revenue per visitor of $3.75 with a standard deviation of around $30. Detecting a 5% change in revenue needs about 409,600 visitors per version, while a 5% change in conversion needs about 121,600. At 20,000 visitors per version, the revenue MDE is 4 × $30 ÷ √20,000 = $0.85, a 23% relative lift, against about 12% for conversion rate at the same traffic.

That doesn't mean you should judge on conversion rate instead. A version can raise conversion and lower revenue. It means revenue tests need bigger changes or more time. See revenue per visitor for how to measure it.

How to choose your MDE

Start from money, not statistics. Ask what lift would be worth the effort of shipping and maintaining the change. A 3% lift on a page that brings in $200,000 a month is worth $6,000 a month. The same lift on a $5,000-a-month page is worth $150. The split test ROI calculator does this sum.

Then check what's realistic. Writing about Microsoft's experiments in Harvard Business Review, Kohavi and Thomke note that most changes to engagement have an effect smaller than 1%. Small businesses making bolder changes can expect bigger swings, but don't plan a test around a 40% lift from a headline rewrite.

Your MDE should sit at or below the smallest lift worth shipping. If your traffic can't get there in six to eight weeks, change the plan: test a bolder idea, test a higher-volume metric, or skip the test.

Outtest uses a related idea. Its Referee only calls a winner that shows a lift of at least 10% by default, adjustable from 2 to 30% in Settings, along with a 90% chance of beating the original and at least 7 days of data.

What MDE doesn't tell you

  • It isn't a guarantee. At 80% power, a real lift exactly equal to your MDE is missed one time in five.
  • It isn't the size of the win. When a test with a large MDE does find a winner, the measured lift tends to be bigger than the true one, because only lucky high readings cross the line. Airbnb researchers call this selection bias the winner's curse and proposed a correction for it (Airbnb Tech Blog).
  • It assumes you don't peek. Checking daily and stopping on the first significant reading breaks the maths. See the peeking problem.

Ways to lower your MDE

  • Run longer, or add traffic. Four times the visitors halves the MDE.
  • Test two versions, not four. Each extra version splits the same traffic further.
  • Count only visitors who could see the change. Kohavi and colleagues show that analyzing only the users who started checkout, for a checkout change, cut the users needed to 6,400 in their example.
  • Measure a metric close to the change, like add-to-cart for a product page layout, as long as you keep revenue as a guardrail.
  • Remove extreme outliers from revenue metrics with a rule agreed before the test.

For the reverse question, how many visitors you need for a given MDE, see A/B test sample size and the sample size calculator.

Questions people ask

What is a good minimum detectable effect?+

One that's smaller than the lift you'd be happy to ship and small enough to be realistic for the change. For big sites, 1 to 5% is common. For small sites, 20 to 30% is often the best the traffic allows, which means only bold changes are worth testing.

Is MDE the same as the lift I'll get?+

No. MDE is a property of the test design, not a prediction. The real effect could be zero, smaller than the MDE, or bigger. If the true lift equals your MDE, the test still misses it about 20% of the time at 80% power.

Should I use relative or absolute MDE?+

Either works if you're clear which one you mean. Absolute MDE is in percentage points (3.0% to 3.5% is 0.5 points). Relative MDE is the percentage change (0.5 ÷ 3.0 is about 17%). Most tools and teams talk in relative terms, so state the baseline too.

How do I lower my MDE?+

Run longer or send more traffic, since quadrupling the sample halves the MDE. Test fewer versions at once, measure a metric closer to the change, and only count visitors who actually reached the page being tested.

Read next

Let Outtest run your split tests

AI agents read your analytics and payments, find where you lose the most money, build the fix and test it. Every test is judged on revenue, not clicks. Plans from $29 a month.