outtest
Get started
What to test

What to A/B test first

Start where your funnel loses the most money, then score ideas on impact, confidence and ease and check you have the traffic. A worked example with real maths.

Updated 28 September 2026 · 6 min read

The short answer

Test first where the most money leaks out of your funnel, usually checkout, pricing or the step with the worst drop-off, not the homepage. Score each idea on impact, confidence and ease (ICE, 1 to 10 each, multiplied), then check the test can finish within about six weeks at your traffic. Steps near the payment often have fewer visitors but higher conversion rates, so they can reach an answer faster than a homepage test.

Free tool: A/B test ideas. No signup.

Test first where your funnel loses the most money. For most online businesses that's a step close to payment, like checkout, cart or the pricing page, not the homepage. Among the ideas for that step, run the one with the best mix of expected impact, evidence and effort, and only if your traffic can give you an answer within about six weeks.

That's three filters: money at stake, a simple score such as ICE, and a traffic check. The traffic check often surprises people, because steps near the payment have fewer visitors but much higher conversion rates, so their tests can finish faster than a homepage test. The worked example below shows all three with the maths.

Step 1. Map the funnel in money

Pull a month of data from your analytics and payment tools and write down how many people reach each step. Here's an example online store (the numbers are made up for illustration):

Step People a month Rate from previous step
Visit the site 60,000
Add to cart 4,800 8%
Start checkout 2,400 50%
Place an order 1,200 50%

At a $50 average order, that's $60,000 a month.

Now work out what a person at each step is worth: revenue divided by the number of people who reached the step.

  • A visitor is worth $60,000 / 60,000 = $1.00
  • A shopper with a cart is worth $60,000 / 4,800 = $12.50
  • A shopper who starts checkout is worth $60,000 / 2,400 = $25.00

So losing one shopper at checkout costs as much as losing 25 visitors at the front door.

The biggest raw drop-off is at the top, where 55,200 visitors leave without adding anything. But most of them were never going to buy, so that number overstates what you can win back. A better way to see the money at stake is to ask what a realistic lift at each step would add.

In a funnel like this, a 10% relative lift at any single step raises revenue by 10%, or $6,000 a month. So the question isn't "which step has the biggest number" but "at which step is a big lift most likely?" That's where the steps with abnormal drop-off and a known cause come in.

This store loses 75% of carts before an order. The Baymard Institute puts the documented average cart abandonment rate at 70.22%, across 50 studies. In Baymard's survey of shoppers who abandoned for reasons other than just browsing, 40% blamed extra costs like shipping and tax, and 18% said the site wanted them to create an account. That's evidence pointing straight at two fixable problems.

Step 2. Score the ideas with ICE

ICE scores each idea from 1 to 10 on three things and multiplies them. ProductPlan credits the method to Sean Ellis.

  • Impact: how much money if it works. Tie this to the money map, not to excitement.
  • Confidence: how strong the evidence is. Session recordings, support tickets, survey answers and research like Baymard's raise it. "I have a feeling" is a 2.
  • Ease: how quickly you can build and launch it. A copy change is a 9. A checkout rebuild is a 3.

Here are four candidate tests for the example store:

Idea Step Plausible lift Money a month if it works Impact Confidence Ease ICE
A. New homepage headline Visit 5% $3,000 3 3 9 81
B. Real review snippets on product pages Visit to cart 10% $6,000 5 4 7 140
C. Show delivery cost in the cart Cart to order 15% $9,000 7 7 6 294
D. Guest checkout Checkout to order 20% $12,000 9 6 4 216

The "plausible lift" column is a judgment call based on the evidence, and that's fine. The point of ICE is to make the judgment explicit so the team can argue about the inputs rather than the conclusion. Some teams average the three scores instead of multiplying; multiplying punishes a weak score harder, which is usually what you want.

On ICE alone, C goes first and D second.

Step 3. Check the traffic

A high score is useless if the test can't finish. Each test's sample size depends on how many people enter it and their conversion rate to an order. Using the standard formula at 95% confidence and 80% power (the /tools/ab-test-sample-size-calculator does this):

Idea Who enters the test Per day Order rate of entrants Lift to detect Needed in total Days
A All visitors 2,000 2% 5% 630,412 315
B All visitors 2,000 2% 10% 161,364 81
C Shoppers with a cart 160 25% 15% 4,386 27
D Checkout starters 80 50% 20% 776 10

Rounded up to whole weeks, C takes 4 weeks and D takes 2. A would take most of a year and B nearly three months, too long to be practical, since traffic, seasons and cookies all drift over that time (/guides/how-long-to-run-an-ab-test explains why tests should end within about four to six weeks).

This is the counterintuitive part. Test D sees only 80 people a day but needs just 776 of them, because half already buy and a 20% lift is a big absolute jump (50% to 60%). Test A sees 2,000 people a day but needs 630,412, because a 5% lift on a 2% rate is a change of 0.1 percentage points.

Step 4. Decide the order

For the example store:

  1. Run C (delivery cost in the cart) now. Highest ICE score, strong evidence, four weeks.
  2. Run D (guest checkout) at the same time on the checkout page. The tests touch different pages, and in Controlled experiments on the web, Kohavi and colleagues report that strong interactions between concurrent experiments are rare in practice.
  3. Don't test A. A headline tweak can't reach an answer at this traffic. Either ship it if it's clearly better copy, or save it for a bold rewrite that changes the promise and could plausibly lift more.
  4. Rework B into something bigger, such as a new product page layout with reviews, delivery dates and returns information together, so the plausible lift justifies the test length.

Where the money usually leaks

Start from your own numbers, but some places show up again and again:

  • Checkout and payment: forced account creation, surprise costs at the last step, few payment methods, a form that breaks on mobile.
  • Pricing pages: which plan is highlighted, monthly versus annual framing, how the plans are named, price points themselves. See /guides/how-to-ab-test-pricing.
  • Trials: whether to offer one, and how long. See /guides/free-trial-vs-no-trial.
  • Cancel flows: a pause option or a downgrade offer can keep customers who would otherwise leave. See /guides/cancel-flow-ab-testing.
  • Landing pages for paid traffic: the promise in the headline and whether the page matches the ad. See /guides/landing-page-ab-testing and /guides/headline-ab-testing.

For more ideas sorted by page and business type, the free /tools/ab-test-ideas list has over 100. To see what a lift would be worth to you per month and per year, use the /tools/split-test-roi-calculator.

Judge each test on money

Whatever you test first, pick the winner on revenue per visitor, not clicks or signups. A cart change that raises checkouts but pushes people to cheaper shipping options, or a pricing change that raises signups on a cheaper plan, can win the easy metric and lose money. Revenue per visitor is noisier than a conversion rate because order values vary, so budget more visitors than the table above suggests. /guides/revenue-per-visitor shows how much more.

This prioritization is the job of two of Outtest's agents. The Analyst reads your analytics and payment data to find where the funnel loses the most money, and the Lead plans which test to run there next. You can do the same by hand once a month with a spreadsheet and the calculators above.

When your traffic is low

If every idea on your list fails the traffic check, you have three options: test bigger changes (a new offer rather than a new button), test closer to the payment where conversion rates are higher, or stop testing that step and make the change based on qualitative evidence. /guides/ab-testing-low-traffic covers each.

Questions people ask

What should I A/B test first on my website?+

The funnel step where you lose the most money and have evidence of a fixable problem, usually checkout, cart or pricing rather than the homepage. Then confirm your traffic can detect a realistic lift at that step within about six weeks. If it can't, pick a bolder change or a busier step.

What is the ICE score?+

ICE stands for impact, confidence and ease. You score each test idea from 1 to 10 on each and multiply the three numbers, so a 7, 6 and 5 scores 210. It's credited to Sean Ellis and is popular because it takes minutes and forces you to state your evidence.

Should I test the homepage first?+

Rarely. Most homepage visitors were never going to buy, so a better headline moves overall revenue only a little, and detecting a small lift on a low conversion rate takes a huge sample. Test the homepage when it's the main entry point for paid traffic and you have a bold change to try.

How many A/B test ideas should I have ready?+

Enough to fill the next two or three months, scored and sorted. Most tests won't win, so a short queue means long gaps between wins. Keep adding ideas from analytics, session recordings, support tickets and customer interviews.

Read next

Let Outtest run your split tests

AI agents read your analytics and payments, find where you lose the most money, build the fix and test it. Every test is judged on revenue, not clicks. Plans from $29 a month.