outtest
Get started
Statistics

How holdout groups prove what your tests actually earned

A holdout keeps a small random share of visitors on the old version after you ship winners, so you can measure what those winners really earned.

Updated 28 September 2026 · 6 min read

The short answer

A holdout group is a small random slice of visitors, often 5 to 10%, who keep seeing the original experience after you ship winning tests. Comparing their revenue per visitor with everyone else's shows what the winners earned together. It's usually less than the sum of the individual test lifts, because of lucky readings, overlapping changes and effects that fade.

Free tool: Revenue per visitor calculator. No signup.

A holdout group is a small, random slice of visitors, usually 5 to 10%, who keep seeing the original experience after you ship the winners of your A/B tests. Everyone else gets the winners. Months later, you compare revenue per visitor between the two groups. The gap is what your winning tests actually earned, together, in the real world.

You need this because the lifts your tests report tend to overstate what you get. Some winners won partly by luck, some changes overlap and eat into each other's gains, and some effects fade once the novelty wears off. A holdout measures all of that at once. It's the difference between "our tests say we added 16%" and "we can show we added 7%".

Why test results overstate what you earn

Lucky winners

When you run many tests and ship only the significant ones, the lifts you record are biased upward. The tests that happened to read high crossed the line, and the ones that read low didn't. Airbnb researchers studied this "statistical selection bias" in the aggregated impact of launched features and proposed a correction for it, calling it the winner's curse (Airbnb Tech Blog).

Overlapping changes

Two tests that each add 1% don't necessarily add 2% together. Disney Streaming points out that the total "may be exactly 2%", or less if the changes cannibalize each other, or more if they compound, and "the experimental data alone cannot tell us" which (Disney Streaming).

Fading effects

A change can win in a two-week test because it's new, then lose its edge. Disney found several features that did well in A/B tests showed "decreasing impact, or inconsistent impact over time".

Disney's own result is the clearest warning. Before holdouts, teams summed individual test results: 10 changes of 1% each would be reported as a 10.46% gain. The holdout showed "the cumulative effect of our product features was much smaller than expected".

Types of holdout

Type Who is held back From what Answers
Per-test holdout A small share, such as 5% One shipped winner Did this winner keep earning after the test?
Rolling holdout A small share Every winner shipped in a period What did all our winners earn together?
Universal holdout A small share, reset each quarter Every product change in a period What did our whole roadmap earn?
Ad holdout People who see no ads A campaign Did the ads cause the sales, or would they have happened anyway?

For most small and mid-sized businesses running website tests, the rolling holdout is the useful one. Meta offers the ad version. You can add a holdout to an A/B test to compare "cost per incremental purchase" rather than raw cost per purchase (Meta Business Help Center).

Worked example

An online store gets 300,000 visitors a month. After a winning test ships, 5% of visitors (15,000 a month) stay on the original version, chosen at random and kept there with a cookie. Over a few months it ships three winners whose tests reported lifts of +6%, +5% and +4% in revenue per visitor.

Compounded, the tests claim 1.06 × 1.05 × 1.04 = 1.158, a 15.8% gain.

In the six months after the third winner ships:

Holdout (5%) Everyone else (95%)
Visitors 90,000 1,710,000
Revenue per visitor $1.52 $1.63

The measured gain is $1.63 ÷ $1.52 − 1 = 7.2%, less than half of what the tests claimed.

How precise is that? Revenue per visitor is noisy. Assume a standard deviation of $14 per visitor (an example figure; use your own). The standard error of the holdout's average is $14 ÷ √90,000 = $0.047, and of the main group's is $14 ÷ √1,710,000 = $0.011. The standard error of the difference is √(0.047² + 0.011²) = $0.048, so the 95% range is $0.11 ± $0.094, or $0.016 to $0.204. In relative terms, the winners added somewhere between about 1% and 13%.

That range is wide, but it still shows the winners made money, probably about half what the test reports suggested. With a 10% holdout, the margin would shrink from ±$0.094 to about ±$0.068.

The value is concrete. The 95% of visitors on the winners bring in $0.11 more each, about 285,000 × $0.11 = $31,350 a month. The holdout costs the gain those 15,000 visitors miss, about $1,650 a month.

How big and how long

A holdout is a test like any other, so size it the same way. Decide the smallest combined lift you'd want to confirm and work out how many visitors the holdout needs to detect it. Disney sized its Hulu holdout to detect a 1% change in hours watched, because "a 1% change is large enough to drive a financially meaningful ad revenue impact". Smaller businesses usually need bigger lifts or longer periods, as the example shows. The minimum detectable effect guide has the formula.

For length, Disney used three months of enrollment and a one-month evaluation period, then reset the holdout with a new set of users each quarter. Longer holdouts measure longer-term effects but cost more in missed gains, in engineering time to keep the old version running, and in the risk of mistakes.

How to run a holdout

  1. Pick the holdout share (5 to 10%) and assign visitors at random, before any test assignment.
  2. Keep them there. Store the assignment so returning visitors stay in the holdout.
  3. Exclude holdout visitors from new tests, so they only ever see the original.
  4. Track revenue per visitor for both groups from your payment tool, not just conversion rate.
  5. Compare at fixed points, such as quarterly, not every day.
  6. Record the sum of test-reported lifts next to the holdout result, so you learn how much your tests tend to overstate.
  7. Reset the holdout on a schedule, and give the old group the winners.

Disney's lessons include two practical warnings. Some changes were too expensive to keep in two versions for months, and as more winners stacked up, code branches multiplied and parts of the holdout experience drifted from what was intended until someone noticed. Keep the holdout simple and check it regularly.

Outtest keeps 5% of visitors on the original after every win, so it can show what winners actually earned rather than what the tests predicted.

Reading the result

Compare the holdout gain with the compounded lifts your tests reported for the same winners.

  • If the two are close, your tests are well calibrated and you can trust their estimates when you forecast.
  • If the holdout gain is positive but much smaller, that's normal. Use the ratio to discount future test results. In the example, 7.2% measured against 15.8% claimed is a ratio of about 0.46, so a test reporting +10% is more likely worth +4 to 5%.
  • If the holdout gain is near zero or negative, check the setup first. Holdout visitors who saw a winner through another device or a cached page make the groups look alike. If the setup is clean, look for a winner that raised first purchases but hurt refunds, renewals or repeat orders.

When a holdout isn't worth it

  • Very low traffic. If 5% of your visitors is a few hundred people a month, the margin of error will be too wide to say anything. Put that effort into bolder tests.
  • Bug fixes, security fixes and legal changes. Never hold these back.
  • Price increases you have to apply to everyone for fairness or legal reasons.
  • Changes that can't sensibly exist in two versions, such as a brand rename.

For deciding whether testing is paying off overall, a holdout is the best evidence you can get. Pair it with the split test ROI guide to turn the measured lift into money against what testing cost.

Questions people ask

What percentage should a holdout group be?+

Usually 5 to 10% of visitors. A smaller holdout costs less in missed gains but gives a wider margin of error, so it needs more months of data. Size it with the same maths as a normal test, for the smallest combined lift you'd want to confirm.

How long should a holdout run?+

Long enough to measure the effects you care about, and short enough that the old experience stays maintainable. Disney Streaming settled on three months of enrollment plus a one-month evaluation period, then reset the group each quarter.

What's the difference between a holdout and a control group?+

A control group is part of one A/B test and ends when the test ends. A holdout keeps seeing the old version after winners ship, sometimes across many tests, so it measures the long-run and combined effect rather than one change in isolation.

Is it unfair to keep customers in a holdout?+

They get the experience everyone had before the tests, not a broken one. Keep holdouts small, time-limited and rotated, and never hold back bug fixes, security fixes or legal changes. If a winner is clearly a big improvement, you can end the holdout early.

Read next

Let Outtest run your split tests

AI agents read your analytics and payments, find where you lose the most money, build the fix and test it. Every test is judged on revenue, not clicks. Plans from $29 a month.