outtest
Get started
Basics

How to run an A/B test, step by step

The ten steps of a trustworthy A/B test, from picking the metric and sizing the sample to checking the split, reading the result and keeping a holdout.

Updated 28 September 2026 · 7 min read

The short answer

Pick one money metric, find the funnel step that leaks most, write a hypothesis with a minimum lift, then calculate the sample size and round the run time up to whole weeks. Split visitors 50/50 by person, check the split after a day or two, and don't stop early. Read the result once at the planned end, ship the winner and keep a small holdout to confirm what it really earned.

Free tool: A/B test duration calculator. No signup.

Running an A/B test comes down to ten steps: choose the metric, find the leak, write a hypothesis, size the test, build the change, split traffic, check the split, wait, read the result once, then ship and keep a holdout. Skipping step four (sizing) and breaking step eight (stopping early when the numbers look good) are the two quickest ways to get a wrong answer.

The rest of this guide walks through each step with a running example: an online store with 1,800 visitors a day to its product pages, 4% of whom buy. The numbers are made up for illustration, but the maths is real.

1. Choose the one metric that decides the winner

Write down the single number that decides the test before anything else. For anything that sells, that should be revenue per visitor, which is total revenue from the people who saw a version divided by the number of people who saw it. Clicks and signups are easy to move without moving money. A pricing page that pushes people to a cheaper plan can raise signups and lower revenue at the same time (/guides/revenue-per-visitor shows the maths).

Also pick two or three guardrail metrics that must not get worse, such as refund rate, page load time or support tickets. You won't pick the winner on them, but you'll stop a version that damages them.

2. Find where the funnel loses the most money

Don't start from a list of ideas. Start from your analytics and ask at which step the most people with money to spend drop out. A checkout that loses 50% of people who start it is usually a better place to test than a homepage headline. /guides/what-to-ab-test-first shows how to score candidate tests by money at stake, evidence and effort.

In our example, session recordings show many shoppers leave the product page without seeing the delivery cost, which appears only at checkout.

3. Write a hypothesis with a minimum lift

A useful hypothesis names the evidence, the change, the audience and the size of effect you need:

"Because recordings show shoppers leave without seeing delivery costs, showing the delivery cost and date on the product page will raise revenue per visitor for product page visitors by at least 15%."

The minimum lift matters because it sets the sample size. Ask "what's the smallest improvement that would be worth the work of shipping and maintaining this?" If a 5% lift wouldn't change your plans, don't size the test to detect 5%.

4. Calculate the sample size and duration

The standard formula for comparing two conversion rates at 95% confidence and 80% power is:

n per version = (1.96 × √(2p̄(1 - p̄)) + 0.8416 × √(p₁(1 - p₁) + p₂(1 - p₂)))² / (p₂ - p₁)²

where p₁ is the current rate, p₂ is the rate after the lift you want to detect, and p̄ is their average.

For the example, p₁ = 4% and a 15% lift gives p₂ = 4.6%:

  • p̄ = 4.3%, so √(2 × 0.043 × 0.957) = 0.2869, times 1.96 = 0.5623
  • √(0.04 × 0.96 + 0.046 × 0.954) = 0.2869, times 0.8416 = 0.2414
  • (0.5623 + 0.2414)² = 0.6459
  • (p₂ - p₁)² = 0.006² = 0.000036
  • n = 0.6459 / 0.000036 ≈ 17,943 visitors per version (rounded up)

This formula covers conversion rates. Revenue per visitor is noisier because order values vary, so if you judge on revenue, treat the result as a floor and add a margin.

That is 35,886 visitors in total. At 1,800 a day, the test needs 20 days, which rounds up to 21 days, or three full weeks. Round up to whole weeks because weekday and weekend shoppers behave differently. Kohavi and colleagues recommend running for at least a week or two and continuing in multiples of a week (Controlled experiments on the web, 2009).

The /tools/ab-test-sample-size-calculator and /tools/ab-test-duration-calculator do this in seconds. If the answer is "four months", test a bigger change or a busier page. /guides/ab-test-sample-size and /guides/how-long-to-run-an-ab-test go deeper.

5. Build the new version and check it

Change one idea at a time. "Show delivery cost and date" is one idea even if it touches two lines of text. "New layout, new copy, new photos and a discount" is four ideas, and when it wins you won't know which one did the work.

Before launch, check:

  • Both versions render correctly on mobile, tablet and desktop, in the main browsers.
  • Tracking fires the same way in both. A tracking bug in one version fakes a result.
  • The new version doesn't load noticeably slower. Speed alone moves revenue. Bing found every 100 milliseconds of delay cost 0.6% of revenue (HBR, 2017).
  • No page flicker, where visitors briefly see the original before the change appears.
  • The copy makes no claim you can't back up. Don't invent reviews, stock levels or discounts to win a test.

6. Split traffic at random, by visitor

Assign each visitor to A or B at random and keep them there on every return visit. Randomizing by page view instead of by person means some people see both versions, which blurs the result.

Split 50/50 unless you have a reason not to. A lopsided split is slower. Kohavi's team notes that a 99/1 split takes about 25 times longer to reach the same power as 50/50 (Controlled experiments on the web). Don't change the split partway through either. Changing allocation mid-test can produce Simpson's paradox, where the combined result points the opposite way from every individual period (Seven Rules of Thumb for Web Site Experimenters, 2014).

7. Check the split after a day or two

Once the test has a few thousand visitors, check that the traffic split matches what you set. If you asked for 50/50 and see 10,250 against 9,750, a chi-square test gives p = 0.0004. Chance almost never produces a gap like that. Something (a redirect, a bot filter, a slow script) is dropping visitors from one version, and the result can't be trusted. The /tools/srm-checker runs the check, and /guides/sample-ratio-mismatch explains the usual causes.

This is not a rare problem. Microsoft researchers found that about 6% of experiments at Microsoft had a sample ratio mismatch (Fabijan et al., KDD 2019).

8. Run to the planned end without stopping early

This is the hardest step. The result will swing in the first days. If you stop the first time it crosses 95%, you'll ship far more false winners than the 5% you signed up for. Evan Miller showed that checking after every visitor can push the real false positive rate to 26.1% (How Not To Run an A/B Test). /guides/peeking-problem-ab-testing shows the full picture.

It's fine to look, and it's fine to stop a version that is clearly broken or losing badly. Just don't declare a winner before the planned end.

9. Read the result once

At the planned end, look at:

  • The lift on your decision metric and its confidence interval, not just the headline number.
  • The p-value or the chance that B beats A (/guides/statistical-significance-explained).
  • The guardrail metrics.

Suppose the example finishes with 17,943 visitors per version, 718 orders on A (4.0%) and 826 on B (4.6%). The z-test gives p ≈ 0.005 and a 95% interval for the lift of roughly +4.5% to +25.5%. The lower end is below the 15% you hoped for, but the whole range is positive, so B wins.

Resist slicing the data into segments until one looks exciting. With 20 segments, the chance that at least one crosses p = 0.05 by luck alone is 1 - 0.95²⁰ = 64%. Treat segment findings as ideas for the next test.

10. Ship, hold out and write it down

Ship the winner to most visitors, but keep a small holdout (say 5%) on the original for a few more weeks. Measured lifts often shrink after launch. Tests that clear the bar include some lucky draws, and a significant result from a low-powered test tends to exaggerate the effect, which Kohavi, Deng and Vermeer discuss as the winner's curse in A/B Testing Intuition Busters. The holdout shows what the change really earns. /guides/holdout-groups covers how to size one.

Then record the test: hypothesis, dates, sample sizes, result, and what you'll try next. A flat or losing test is still a finding. It stops the next person from running the same idea.

The checklist

Step Done when
Metric One decision metric and guardrails written down
Leak You can say which funnel step you're fixing and why
Hypothesis Evidence, change, audience and minimum lift written down
Size Sample size and end date calculated, rounded to whole weeks
Build Both versions checked on real devices, tracking verified
Split Random by visitor, sticky, 50/50
Health Split checked for mismatch after a day or two
Wait Test ran to the planned end
Read Lift, interval, p-value and guardrails reviewed once
Ship Winner live, holdout running, test logged

Outtest runs these ten steps as a loop. Its Analyst finds the leak, the Lead picks the test, the Builder writes the new version without inventing claims, and the Referee calls the result with plain maths. Changes to checkout, pricing and cancel flows always wait for your approval, and 5% of visitors keep seeing the original after a win so you can see what the winner actually earned.

Questions people ask

What are the steps of an A/B test?+

Choose the metric, find where the funnel leaks, write a hypothesis, calculate sample size and duration, build and check the new version, split traffic at random by visitor, check the split is healthy, run to the planned end without stopping early, analyze once, then ship and keep a holdout. Skipping the sizing step is the most common reason tests end with nothing learned.

How many versions should an A/B test have?+

Two is best for most sites: the original and one change. Every extra version takes a share of the traffic and adds another chance of a false win, so a four-version test needs far more visitors to reach the same confidence.

Should I split traffic 50/50?+

Yes, unless you have a reason not to. A 50/50 split reaches a result fastest. Kohavi and colleagues note that a 99/1 split takes about 25 times longer to reach the same power as 50/50.

What should I do if my A/B test shows no difference?+

Keep the original and record the result. A flat result tells you the change doesn't matter at the size you could detect, which is useful: it stops you shipping work that does nothing. Next time, test a bigger change or a different step of the funnel.

Read next

Let Outtest run your split tests

AI agents read your analytics and payments, find where you lose the most money, build the fix and test it. Every test is judged on revenue, not clicks. Plans from $29 a month.