outtest
Get started
Basics

A/B testing vs multivariate testing

A/B tests compare whole versions of a page; multivariate tests change several elements at once in every combination. How they differ and when each is worth it.

Updated 28 September 2026 · 6 min read

The short answer

An A/B test compares the original page with one or more complete alternatives, while a multivariate test changes several elements at once (say two headlines, two images and two buttons, making 8 combinations) to measure each element's effect and how they interact. Use A/B tests for most decisions. Multivariate tests pay off only on high-traffic pages, because picking the best of 8 combinations takes about 4 times the visitors of a simple A/B test.

Free tool: A/B test sample size calculator. No signup.

An A/B test compares whole versions of a page: the original against one alternative, or against several (that's an A/B/n test). A multivariate test (MVT) changes several elements of the page at the same time and shows visitors every combination, so you can estimate how much each element matters and whether elements help or hurt each other.

For most sites the answer is to run A/B tests. Multivariate tests make sense only when a page gets heavy traffic, you have several independent changes ready at once, and you have a real question about how those changes interact. The reason is traffic: depending on what you want to learn, a multivariate test can cost the same as one A/B test or four times as much.

The difference at a glance

A/B test A/B/n test Multivariate test
What changes One idea, or a whole redesign Several complete alternatives Several elements, each with 2 or more options
Groups 2 3 or more Every combination (2 × 2 × 2 = 8)
Answers Is B better than A? Which of these versions is best? Which elements matter, and do they interact?
Traffic needed Lowest Grows with each version Low for main effects, high for picking a combination or measuring interactions
Setup Simple Simple All variants must be ready at launch
Best for Most decisions, low and medium traffic Comparing a few distinct ideas High-traffic pages with several independent changes

How a multivariate test works

Say a landing page has three elements you want to change: the headline (current or new), the hero image (current or new) and the button text (current or new). A full factorial multivariate test serves all 2 × 2 × 2 = 8 combinations, each to an eighth of visitors.

Because every element appears in every pairing, you can analyze the results two ways:

  • Main effects: compare everyone who saw the new headline (4 combinations) with everyone who saw the old one (the other 4). Do the same for the image and the button.
  • Interactions: check whether the new headline works better with the new image than with the old one.

When there are too many combinations, testers run a fraction of them in a planned pattern, called a fractional factorial or Taguchi design. Ronny Kohavi and colleagues describe these designs in Controlled experiments on the web, where they use the term multivariable testing.

The traffic maths, worked through

Take a page with a 3% conversion rate and a goal of detecting a 20% relative lift (3.0% to 3.6%) at 95% confidence and 80% power. The standard sample size formula gives 13,914 visitors per group. The /tools/ab-test-sample-size-calculator returns the same number.

A simple A/B test needs 2 × 13,914 = 27,828 visitors.

Now the 2 × 2 × 2 multivariate test, three ways:

  1. Main effects only. Each main effect compares 4 combinations against the other 4, so each side of the comparison gets half the traffic, exactly like an A/B test. With 27,828 visitors you can estimate all three main effects at the same power, as long as the elements don't interact strongly. Testing the three changes one after another as A/B tests would take 3 × 27,828 = 83,484 visitors. This is the efficiency argument for MVT: Kohavi's team gives the example of five changes that would take five months as sequential four-week A/B tests, or one month as a single multivariate test with the same power.

  2. Picking the best combination. If you want to know which of the 8 combinations wins, each one needs its own 13,914 visitors: 8 × 13,914 = 111,312 visitors, four times the A/B test. You're also making 7 comparisons against the original. At a 0.05 bar each, the chance that at least one "wins" by luck is 1 - 0.95⁷ = 30%, so you'd need a stricter bar and even more traffic.

  3. Measuring an interaction. An interaction is a difference of differences: (new headline lift with new image) minus (new headline lift with old image). Each of those lifts comes from a quarter of the traffic, so the estimate is twice as noisy as a main effect, and detecting an interaction of the same size takes about four times the visitors.

So "multivariate tests need more traffic" is only true for questions 2 and 3. It's still the practical rule, because most people running an MVT want to pick the winning combination.

Do interactions actually matter?

Sometimes. Kohavi and colleagues give an example of two changes on a product page, a larger product image and more product detail, that might each raise sales alone but together push the buy button below the fold and lower them. That's an interaction, and it's the kind you should catch in planning rather than in a test.

In practice, strong interactions are uncommon. The same 2009 paper says "strong interactions are rare in practice" and recommends that your first experiment be an A/B test. In Seven Rules of Thumb for Web Site Experimenters (2014), the authors, who ran experiments at Microsoft and LinkedIn, say they usually find it more useful to run simple designs that test one or two variables at a time.

If interactions are rare, you gain little from an MVT's ability to measure them, and you pay for it in complexity.

The other costs of multivariate tests

Traffic isn't the only cost. The 2009 paper lists three limitations:

  • Some combinations may be bad experiences that nobody would ship, but a full factorial shows them to real visitors anyway.
  • Analysis is harder. You get many comparisons plus interactions, which makes deciding what to ship more complex.
  • Every variant must be ready before launch. If one element is delayed, the whole test waits. With A/B tests you run whichever change is ready first.

There's also a metric problem. If you judge tests on revenue per visitor rather than a click rate, as you should for anything that sells, each group's number is noisier because order values vary. That raises the sample size for every cell and makes a many-cell design even slower. /guides/revenue-per-visitor explains why revenue is still the right thing to measure.

Where A/B/n tests fit

An A/B/n test compares the original with several complete alternatives: four headlines, say. It's useful when you have a few distinct ideas and no reason to combine them.

It has the same multiple comparison problem as picking an MVT combination. With four challengers each tested at 0.05, the chance of at least one false win is 1 - 0.95⁴ = 18.5%. The simple fix (a Bonferroni correction) is to divide the bar by the number of comparisons: 0.05 / 4 = 0.0125. At that stricter bar, each group needs 19,768 visitors to detect a 20% lift on a 3% rate, so five groups need 98,840 visitors. The plain A/B test needed 27,828.

That's why "test four headlines at once" is often slower than testing the strongest headline first and the runner-up next.

Which one should you use?

Your situation Use
Under about 1,000 visitors a day to the page A/B test with one bold change (see /guides/ab-testing-low-traffic)
One clear idea, or a full redesign A/B test
Three or four distinct ideas, decent traffic A/B/n test with a corrected significance bar, or sequential A/B tests
Several small, independent changes on a high-traffic page, all ready now Multivariate test, analyzed for main effects
A specific worry that two changes clash Plan them so they don't, or run a 2 × 2 test and budget four times the traffic for the interaction

A useful habit is to write down the decision the test will drive before choosing the design. If the decision is "ship the new headline or not", you need an A/B test. If it's "which of these 8 combinations do we ship", you need the traffic for 8 groups.

Outtest tests headlines, button placement and page layout as split tests judged on revenue per visitor, so the traffic maths above applies there too. Fewer groups and bigger changes reach a clear answer sooner.

Questions people ask

What is the difference between A/B testing and multivariate testing?+

An A/B test compares complete versions of a page, usually the original against one change. A multivariate test changes several elements at once and shows visitors every combination, so it can estimate the effect of each element and whether elements work better or worse together. Multivariate tests need more traffic when you want to compare individual combinations or measure interactions.

Does multivariate testing need more traffic than A/B testing?+

It depends on the question. If you analyze a full factorial test for each element's main effect, it can use the same traffic as one A/B test. If you want to pick the single best combination out of 8, each combination needs its own sample, which takes about 4 times as many visitors as a two-version A/B test.

When should I use multivariate testing?+

When a page gets tens of thousands of visitors a week, you have several independent changes ready at the same time, and you suspect they affect each other. Kohavi and colleagues recommend making your first experiment an A/B test and note that strong interactions between elements are rare in practice.

What is an A/B/n test?+

An A/B/n test compares the original with several complete alternatives, for example four different headlines. It sits between A/B and multivariate testing. Each extra version takes a share of traffic and adds another chance of a false win, so you need a stricter significance bar or more visitors.

Read next

Let Outtest run your split tests

AI agents read your analytics and payments, find where you lose the most money, build the fix and test it. Every test is judged on revenue, not clicks. Plans from $29 a month.