Split testing for marketing agencies
How marketing agencies should run split tests for clients: which clients have the traffic, which tests to run first, and how to report wins on revenue.
Updated 28 September 2026 · 4 min read
Sort clients by traffic before you promise tests. A client with 3,000 visitors a month can only detect changes of around 90% in four weeks, while one with 60,000 can detect about 17%. Run the same short playbook on each client that qualifies (ad landing page, offer, headline, ad copy), and report every result as revenue per visitor from the client's payment data.
An agency's testing problem is different from an in-house team's. You have many sites, each with its own traffic, stack and approval chain, and clients who judge you on revenue. The work is picking which clients can support tests, running a repeatable set of first tests, and reporting results in money.
Why testing is different for agencies
Traffic varies a lot between clients. A test plan that works for your biggest client will stall on your smallest. Before you sell testing, work out what each client's traffic can detect.
Clients pay for revenue. A test that raised click-through rate while revenue stayed flat is hard to defend on a monthly call. Getting read-only access to the client's payment data early lets you report every test in revenue per visitor. See revenue per visitor.
Approval chains slow everything down. Pricing, checkout and anything that makes a claim about the product usually needs sign-off from the client, sometimes from their legal team. Agree up front who approves what.
You need to show results months later. Clients ask "what did testing actually earn us?" A small group of visitors who keep seeing the original after a win gives you a real answer instead of an estimate.
Which clients can support tests
Here's a made-up client list, using the standard settings (95% confidence, 80% power) and a four-week test:
| Client | Visitors a month to the tested page | Conversion rate | Smallest lift detectable in 4 weeks |
|---|---|---|---|
| A, local service | 3,000 | 2% | about 89% |
| B, SaaS | 15,000 | 3% | about 29% |
| C, DTC store | 60,000 | 2% | about 17% |
Client A can only detect a change that nearly doubles conversions. Tests there should be large (a new offer, a new page) or skipped in favor of fixes you can justify without a test. Client B can test bold changes to pricing and offers. Client C can test most things on this page.
These numbers use conversion rate. A test judged on revenue per visitor needs somewhat more traffic, because order values vary. Run each client through the minimum detectable effect calculator, and read the low traffic guide before you pitch testing to small clients.
The tests to run first for a new client
A short, repeatable playbook keeps quality consistent across clients. Adapt it per client, but start here.
1. A landing page matched to the top ad
Build a page around the promise of the client's biggest paid campaign and test it against the current destination. Measure revenue per ad visitor. See the landing page guide.
2. The offer shown first
On a pricing page, default to annual billing. On a store, show the bundle first. Measure revenue per visitor. The pricing guide covers the setup.
3. A headline that names the buyer
Replace a brand slogan with a headline naming who the product is for and what it does. Measure revenue per visitor.
4. Ad copy angle
In Meta ads, test copy that leads with the customer's problem against the current copy. Measure revenue per ad visitor from the client's payment data, not cost per click. See Meta ads A/B testing.
5. A pause option in the cancel flow
For subscription clients, offer a pause before the cancel button. Measure revenue retained per customer who starts cancelling.
6. A search page-group test
For content-heavy clients, change titles or intros on one group of similar pages and compare them with a matched group you left alone. Measure organic clicks and revenue from organic visitors. See SEO split testing.
Common mistakes
- Promising a percentage lift in the pitch.
- Starting tests before you have payment data, then reporting clicks because that's all you have.
- Running small copy tests on clients like A above, then explaining a string of draws.
- Shipping a pricing or checkout change without written client approval.
- Running tests on the same funnel in parallel, so no one can tell which change did what.
- Reporting only winners. Draws and losers show the client you aren't guessing.
How Outtest fits
Outtest connects read-only to a client's payment and analytics tools, finds where their funnel loses the most money, and builds, launches and monitors tests like these. It connects to most common stacks, including Stripe, Shopify, Paddle, Whop, Chargebee and RevenueCat for payments, and Google Analytics, PostHog and Mixpanel for analytics. Every website, pricing and ad test is judged on revenue per visitor.
The Referee agent calls winners with plain maths. By default a version needs at least 7 days, a 90% chance of beating the original and a lift of at least 10%, and a test that misses those bars by day 42 is a draw. Pricing, checkout and cancel flow changes always wait for approval. After a win, 5% of visitors keep seeing the original, which gives you a measured number for the client report.
Every Outtest plan covers one site, so each client site needs its own plan, from Starter at $29 a month for 3 tests to Scale at $499 for unlimited tests.
Questions people ask
How do agencies decide which clients to run A/B tests for?+
Check each client's monthly visitors to the page you'd test and its conversion rate, then work out the smallest lift that traffic can detect in four weeks. If the answer is above 30 to 40%, only a big change will show up, and small copy tests will end as draws.
How should an agency report A/B test results to clients?+
On revenue per visitor from the client's payment tool, with the lift, the chance the new version beats the original and the dates. Leave click-through rate out of the headline number. Report draws too, since a draw means the original stays and nothing was lost.
Should an agency promise a conversion lift in the pitch?+
No. Nobody knows in advance which tests will win, and many end in draws. Promise a process instead: how many tests a month, how winners are called, and how results are reported.
Who approves changes when an agency tests for a client?+
The client, for anything touching price, checkout, cancellation or claims about the product. Agree this in writing before the first test, including who signs off and how fast.
Read next
Let Outtest run your split tests
AI agents read your analytics and payments, find where you lose the most money, build the fix and test it. Every test is judged on revenue, not clicks. Plans from $29 a month.