outtest
Get started
Statistics

What is sample ratio mismatch (SRM) and how do you check for it?

Sample ratio mismatch means your A/B test's traffic split doesn't match what you set. How to check with a chi-square test, worked examples, causes and fixes.

Updated 28 September 2026 · 6 min read

The short answer

Sample ratio mismatch (SRM) happens when the visitor counts in an A/B test's versions differ from the split you configured by more than chance can explain, for example 10,250 against 9,750 on a 50/50 test. Check it with a chi-square goodness-of-fit test on the visitor counts; a p-value below a strict bar such as 0.01 or 0.001 means the test is broken and its result shouldn't be trusted. It's common: about 6% of experiments at Microsoft had one.

Free tool: Sample ratio mismatch (SRM) checker. No signup.

A sample ratio mismatch (SRM) is when the number of visitors in each version of an A/B test doesn't match the split you set, by more than chance can explain. If you configured 50/50 and ended with 10,250 visitors in A and 9,750 in B, that's not bad luck. Something in the setup is removing visitors from one version or adding them to the other, and the test result can't be trusted.

You check for it with a chi-square goodness-of-fit test on the visitor counts. It takes a minute, needs no conversion data, and catches a large share of broken tests. The /tools/srm-checker runs it for any split, including uneven ones.

Why a broken split ruins the result

Random assignment is what makes an A/B test work. The two groups should be alike in everything except the change. When one version loses visitors, the visitors it loses aren't random. They're the ones on slow phones, or the ones who hit a failed redirect, or the ones a bot filter misread. The remaining groups are no longer comparable, and the lift you measure mixes the effect of your change with the effect of who went missing.

Lukas Vermeer's SRM Checker FAQ describes a real 50/50 test that recorded 9,463 users in the control and 7,681 in the treatment. The treatment took about five seconds longer to load, and many of the people who didn't wait were never counted. The test showed the treatment converting slightly better. Counting the missing users as non-buyers, it could have been about 18% worse.

How common it is

More common than most teams assume:

  • About 6% of experiments at Microsoft showed an SRM, according to Fabijan et al., KDD 2019. The authors note that a product running ten thousand experiments a year can expect at least one SRM a day.
  • At LinkedIn, about 10% of triggered experiments (analyses limited to the users a change actually affected) used to suffer from this kind of bias (Chen, Liu and Xu, 2018).

Both companies had spent years fixing the root causes in their platforms, so there's little reason to expect a smaller setup built on page scripts to do better.

How to check, with the maths

The chi-square goodness-of-fit test compares the counts you observed with the counts your planned split predicts:

χ² = Σ (observed - expected)² / expected

summed over every version, with degrees of freedom equal to the number of versions minus one. For a two-version test that's 1 degree of freedom. You then read the p-value from the chi-square distribution (any spreadsheet does this with CHISQ.DIST.RT, and the /tools/srm-checker does it for you).

Example 1. A 50/50 test

The test ends with 10,250 visitors in A and 9,750 in B. Total: 20,000, so the expected count is 10,000 each.

  • A: (10,250 - 10,000)² / 10,000 = 62,500 / 10,000 = 6.25
  • B: (9,750 - 10,000)² / 10,000 = 6.25
  • χ² = 12.5, with 1 degree of freedom
  • p = 0.0004

A 51.25/48.75 split looks harmless, but chance produces a gap this large about 4 times in 10,000. That's an SRM.

Compare 10,100 against 9,900. Each version is off by 100, so χ² = 1 + 1 = 2.0 and p = 0.157. That's ordinary noise.

Example 2. An uneven split

Uneven splits are common: a 90/10 test that exposes a risky change to few people, or a holdout that keeps 10% on the original. With 20,000 visitors, the expected counts are 18,000 and 2,000. Suppose you observe 17,850 and 2,150.

  • Large group: (17,850 - 18,000)² / 18,000 = 22,500 / 18,000 = 1.25
  • Small group: (2,150 - 2,000)² / 2,000 = 22,500 / 2,000 = 11.25
  • χ² = 12.5, p = 0.0004

An SRM again. Notice the small group carries most of the signal. An extra 150 visitors barely registers against 18,000 but is 7.5% too many against 2,000. That's why eyeballing percentages fails on uneven splits, because 89.25/10.75 looks close to 90/10.

Example 3. A tiny gap at scale

Kohavi and Thomke's HBR article gives an example from Microsoft: 821,588 users against 815,482 on a 50/50 test, a 50.2/49.8 split.

  • Total 1,637,070, expected 818,535 each
  • Each version is off by 3,053, so χ² = 2 × 3,053² / 818,535 = 22.8
  • p ≈ 0.0000018, about 1 in 550,000

The article puts it as less than one in 500,000. A difference of 0.2 percentage points is a serious problem at this scale.

Which threshold to use

The usual 0.05 bar is too loose for a health check you run on every test, since it would raise a false alarm on 1 test in 20. Stricter bars are standard:

  • The SRM Checker browser extension built by Vermeer and colleagues flagged a mismatch below p = 0.01.
  • Outtest's free /tools/srm-checker flags one below p = 0.001.

At p = 0.001, a healthy test raises a false alarm about once in a thousand checks. Anything below that deserves investigation, and anything far below it (like the 0.0004 above) almost certainly has a cause you can find.

What causes it

Fabijan and colleagues group causes by the stage where they happen: assignment, running the experiment, processing the logs, and analysis. The ones that show up most on ordinary websites:

  • Redirects. If only the new version redirects to another URL, some visitors drop out during the redirect. The paper's advice is to redirect both versions if one needs it.
  • A slower version. If B loads more slowly, more visitors leave before its tracking fires, so B looks smaller. This was the cause in Vermeer's example above.
  • Bot filtering. In one case in the Fabijan paper, real users engaged so heavily with a new feature that the bot filter classified some of them as bots and removed them from the treatment. Once that was accounted for, the result flipped.
  • Caching. Microsoft once traced an imbalance in an A/A test on its homepage to page caching. The cache key didn't record whether a visitor had a cookie, so tracking was switched on or off for whole batches of visitors (Online Experimentation at Microsoft).
  • Lost cookies. The same Microsoft paper describes a support site where some new users assigned to the treatment were never issued the cookie, so they fell out of it.
  • Analysis filters. Filtering to "users who saw the new widget" when only one version has the widget, or applying different exclusions to each version.
  • Changing the split mid-test. Moving from 90/10 to 50/50 partway through means the totals no longer match either split.

How to find the cause

The Fabijan paper and Vermeer's FAQ list rules of thumb for diagnosis:

  • Split the check by segment: browser, device, country, new against returning. An SRM that appears only on one browser points to a compatibility problem there.
  • Split it by day. An SRM that is strongest on day one and fades suggests caching or a delayed start.
  • Look at page speed and errors per version. A big performance gap often explains the missing visitors.
  • Check whether the SRM appears in the full data or only after filtering. If only after, the filter is the problem.
  • Run an A/A test. If identical versions show an SRM, the problem is in your setup, not your change.

What to do when you find one

  1. Don't ship the result, however good it looks. The bias can run either way.
  2. Find the cause using the checks above.
  3. Fix it and rerun the test from scratch.

If you need a rough answer before rerunning, Vermeer's FAQ describes extreme value bounds. Assume the missing visitors all behaved in the worst way (none converted) and then the best way, and see whether the conclusion survives both. If it flips between the two, the test can't tell you anything yet.

When to check

Check once the test has a few thousand visitors, again at the planned end, and whenever a result surprises you. A check on day one or two catches setup bugs before they waste weeks of traffic.

It also belongs on your pre-read checklist alongside sample size and run time. /guides/ab-testing-mistakes covers the rest of that list, and /guides/how-to-run-an-ab-test shows where the check fits in the process.

Questions people ask

What is sample ratio mismatch in A/B testing?+

It's when the share of visitors in each version doesn't match the split you set, beyond what random chance explains. A 50/50 test that ends with 10,250 visitors in A and 9,750 in B has a mismatch, because a chi-square test gives p = 0.0004. It usually means something is dropping or adding visitors in one version.

How do you test for sample ratio mismatch?+

Run a chi-square goodness-of-fit test on the visitor counts. For each version, take (observed - expected)² / expected, add them up, and look up the p-value with degrees of freedom equal to the number of versions minus one. A very small p-value, below 0.01 or 0.001, signals a mismatch.

What causes sample ratio mismatch?+

Common causes are redirects that fail for some visitors in one version, a version that loads slower so its tracking fires less often, bots filtered differently between versions, caching, and analysis filters that treat versions differently. The fix depends on the cause, so investigate before rerunning.

Can I still use the results of a test with SRM?+

Not as they stand. The missing or extra visitors aren't random, so the comparison is biased, and the bias can point either way. Find and fix the cause, then rerun. If you must estimate something, bound the result by assuming the best and worst outcomes for the missing visitors.

Is a small imbalance like 50.4/49.6 a problem?+

It depends on the sample size. On 2,000 visitors it's normal noise. On 1.6 million visitors, a 50.2/49.8 split has less than a one in 500,000 chance of happening by luck and signals a real problem.

Read next

Let Outtest run your split tests

AI agents read your analytics and payments, find where you lose the most money, build the fix and test it. Every test is judged on revenue, not clicks. Plans from $29 a month.