outtest
Get started
For ai app founders

A/B testing for AI startups

How AI founders should A/B test credit and usage pricing, free credits versus trials and pack sizes, and why heavy users make revenue tests need more traffic.

Updated 28 September 2026 · 4 min read

The short answer

Test pricing before anything else: how many free credits to give, credits versus a time-limited trial, and pack sizes. Judge each test on revenue per visitor, then subtract what the free usage cost you, because every free user burns inference. Heavy buyers make revenue noisy, so a revenue test can need about three times the traffic of a conversion rate test.

Free tool: Revenue per visitor calculator. No signup.

Most software can give away a free plan for next to nothing. An AI product can't, because every prompt, image or minute of audio runs up an inference bill. That makes the free offer and the pricing model the most important things to test, and it means a test can win on revenue and still lose money.

Start with pricing, judge on revenue per visitor, and check the cost of free usage on every winner.

Why testing is different for AI startups

Free usage has a cost. A generous free tier can bring in more paying customers and cost more in inference than they pay. Revenue per visitor tells you which version brought in more money. Subtract the free usage cost in each group before you ship.

Pricing is unsettled. Credits, usage billing, seats and flat plans with limits are all common, and there's no settled default. Every change is a chance to test instead of guess. See how to A/B test pricing.

Revenue is lumpy. With credit packs and top-ups, one customer spends $10 and another spends $400. That spread makes revenue per visitor noisy, so tests need more traffic than the conversion rate alone suggests.

Traffic arrives in spikes. A launch or a viral post brings a flood of curious visitors who behave differently from people who found you by searching. Keep big spikes out of your tests, or let a test run long enough that one spike is a small share of it.

The tests to run first

1. The number of free credits

Test your current free credits against double, or against half. Measure revenue per visitor over 30 days, then subtract the inference cost of free usage in each group.

2. Free credits versus a time-limited trial

Give half of new visitors a fixed number of credits and half a 7-day trial with a usage cap. Measure revenue per visitor after the last trial ends, net of free usage cost. The free trial vs no trial guide covers the timing.

3. A bigger credit pack

Add a larger pack with a lower price per credit above your current packs. Measure revenue per visitor and the pack mix.

4. Monthly plan with credits versus pay as you go

Test a subscription that includes a monthly credit allowance against one-off packs. Measure revenue per visitor over 60 days, since subscriptions pay again and packs may not.

5. Price in outputs, not credits

Show "about 40 images" or "about 3 hours of audio" next to each plan instead of "400 credits". Measure revenue per visitor.

6. Card required for free credits

Ask for a card before giving free credits. Expect fewer signups and less abuse of the free tier. Judge on revenue per visitor, net of free usage cost.

7. A real output on the homepage

Show an actual result the product produced above the fold, next to the input that made it. Measure revenue per visitor. More in the landing page guide.

How much traffic you need

A made-up example. Your pricing page gets 45,000 visitors a month, about 1,500 a day. 2% of visitors pay within 30 days, and payers spend $40 on average, but with a wide spread: most buy a $10 or $20 pack and a few spend hundreds.

With standard settings (95% confidence, 80% power), detecting a 20% lift:

What you judge on Visitors per version Days at 1,500 a day
Conversion rate 21,109 about 28
Revenue per visitor 63,380 about 85

The revenue test needs about three times the traffic, because a handful of big spenders in one group can swing the average. That's still the right test, since the point of a pricing change is revenue. To cope, test bold changes that could move revenue by 25% or more, and check whether a winner still wins with the top few spenders removed. Try your own numbers in the revenue per visitor calculator.

Common mistakes

  • Judging a free credits test on signups. Signups don't tell you whether the credits paid for themselves.
  • Ignoring inference cost on the winning version.
  • Letting a viral spike make up half of a test's traffic.
  • Calling a pricing test after a week when a pack lasts a month.
  • Showing existing customers a new price. Test on new visitors only.

How Outtest fits

Outtest connects read-only to your payment tool (Stripe, Polar, Paddle, Lemon Squeezy, Dodo Payments, Creem and others) and your analytics, and tests pricing pages and pricing copy, trial versus no trial and trial length, onboarding copy and landing pages. Each test is judged on revenue per visitor from your payment data. Outtest does not see your inference costs, so the margin check on a free credits test is yours to run.

Pricing and checkout changes always wait for your approval, and bigger changes like trial length arrive as a GitHub pull request you merge. A version wins after at least 7 days, a 90% chance of beating the original and a lift of at least 10% (defaults you can change). If you want AI assistants to recommend your product, Outtest also runs AI SEO page-group tests, covered in AI SEO testing.

Questions people ask

Should an AI app give free credits or a free trial?+

Test it. Free credits let people try the product at a fixed cost to you, while a time-limited trial lets heavy users run up inference bills. Split new visitors between the two and compare revenue per visitor over 30 days, minus the cost of free usage in each group.

How many free credits should an AI app give?+

Enough for one clearly successful result and not much more, but the right number is specific to your product. Test your current amount against half and double, and judge on revenue per visitor net of the inference cost of the free credits.

Why do AI pricing tests need so much traffic?+

Revenue per paying user varies a lot. Some buy the smallest pack, others top up many times. That spread makes revenue per visitor noisy, so a test needs more visitors to separate real differences from a few heavy buyers landing in one group.

What should an AI startup A/B test first?+

Pricing. AI products have no settled pricing model yet, and a change to free credits, pack sizes or plan structure moves revenue far more than a headline edit. After pricing, test the first screen after signup, which decides whether people see a good result before their credits run out.

Read next

Let Outtest run your split tests

AI agents read your analytics and payments, find where you lose the most money, build the fix and test it. Every test is judged on revenue, not clicks. Plans from $29 a month.