outtest
Get started
What to test

How to A/B test your onboarding

Pick the activation event that predicts paying, track time to first value, and decide onboarding A/B tests on revenue measured 30 to 90 days later.

Updated 28 September 2026 · 6 min read

The short answer

Assign new signups at random to onboarding versions, track time to first value and activation as early signals, and pick the winner on revenue per signup 30 to 90 days later. Activation can move much more than revenue, so a version that lifts activation by 30% may lift revenue by 4% or not at all. Define activation from your own data as the early action that best predicts paying.

Free tool: A/B test duration calculator. No signup.

To A/B test onboarding, assign each new signup at random to the current onboarding or a new version, and keep them in it. Track two early signals: time to first value, meaning how long it takes a new user to get something real out of the product, and activation, meaning the share who complete the early action that best predicts paying. Then decide the winner on revenue per signup 30 to 90 days later.

The early signals exist because revenue is slow. They are not the goal. Onboarding changes can make the activation event easier to reach without making people more successful, so activation jumps and revenue barely moves. A test that only reports activation will ship those changes. A test judged on revenue won't.

The three numbers

Number What it is When you can read it
Time to first value Median time from signup to the first real result Hours to days
Activation rate Share of signups who do the key action within N days Days
Revenue per signup Payments from each version's signups within a fixed window, divided by signups 30 to 90 days

If the onboarding change happens before signup, such as a shorter signup form, use revenue per visitor instead, since the change affects how many people sign up in the first place.

The first session carries a lot of weight. RevenueCat's State of Subscription Apps 2026 found that 55.4% of all 3-day trial cancellations in apps happened on the day the trial started. Many people decide almost immediately.

Find your activation event

Your activation event should come from your own data, not a template. Here is a worked example with made-up numbers. An invoicing tool looks at 2,000 signups from last quarter.

Candidate action in the first 3 days Did it Paid Didn't Paid
Sent a first invoice 700 147 (21%) 1,300 52 (4%)
Completed their profile 1,200 132 (11%) 800 67 (8.4%)

Sending an invoice separates payers from non-payers far better than completing a profile. Its paid rates are 21% against 4%, a gap of 17 points, compared with 11% against 8.4% for the profile. So "sent a first invoice within 3 days" becomes the activation event, and time to first value is the time from signup to that first invoice.

This is a correlation. People who were always going to pay may simply send invoices sooner. That is why activation is a signal to watch and revenue is still the decision.

Randomize the right unit

  • Assign at signup, and keep the assignment for the life of the test.
  • Only include new signups. Existing users have already been through onboarding.
  • If people sign up as teams, randomize by account or workspace, not by person.
  • Run the test over full weeks. People who sign up on Monday for work may behave differently from people who sign up on Saturday.

A worked example

This is a made-up example. The invoicing tool costs $49 a month after a 14-day trial. Version A is the current six-step setup. Version B cuts it to three steps and pre-fills a sample invoice the user can edit and send. Each version gets 1,000 signups.

A: six-step setup B: three steps and a sample invoice
Signups 1,000 1,000
Activated (sent an invoice within 3 days) 350 (35%) 460 (46%)
Median time to first invoice 26 minutes 9 minutes
Paid within 30 days 80 (8.0%) 86 (8.6%)
Average payments per customer in 90 days 2.50 2.42
Revenue per signup at 90 days $9.80 $10.20

The activation result is clear. The standard error of the difference is the square root of (0.35 × 0.65 ÷ 1,000 + 0.46 × 0.54 ÷ 1,000), about 2.2 percentage points. The gap is 11 points, about five standard errors, so B almost certainly activates more people. That is a 31% relative lift.

The paid result is not clear. The standard error for 8.0% against 8.6% is the square root of (0.08 × 0.92 ÷ 1,000 + 0.086 × 0.914 ÷ 1,000), about 1.2 points. The gap is 0.6 points, which gives B about a 69% chance of converting more people to paid. Revenue per signup is 80 × $49 × 2.50 = $9,800 for A and 86 × $49 × 2.42 = $10,198 for B, a 4% difference. The sample is far too small to call that.

So a 31% lift in activation turned into a 4% lift in revenue that may be noise. In this example, some people in B sent the sample invoice without setting up their real business. The change probably helps, but not by anything like 31%.

What to do with a result like this:

  1. Keep B if you have no reason to think it hurts, since activation and time to first value both improved and revenue was no worse.
  2. Keep a holdout group on A so you can check revenue again at six months.
  3. Next time, plan the sample size on revenue, not activation. The sample size guide shows how, and the test duration calculator tells you how long enrollment will take.

Why activation can mislead

One of the rules of thumb in Kohavi and colleagues' paper on online experiments is that reducing abandonment is hard, while shifting where people click is easy. Onboarding is full of easy shifts:

  • Removing steps can raise activation because the event got easier, not because users got more value.
  • A checklist can raise checklist completion without changing whether people come back.
  • A sample project can count as "first value" when the user hasn't touched their own work yet.
  • A new flow can get a short-term bump from novelty that fades.

Each of these lifts the early signal. Only revenue, or retention that you know turns into revenue, tells you whether the change worked.

Over time, you can learn how much your activation lifts carry through. If five past tests show revenue moved by roughly a third of the activation lift, you can use that ratio to decide which new tests are worth a long revenue read.

What to test in onboarding

  1. Fewer signup fields. Measure revenue per visitor, since this happens before signup.
  2. A welcome question ("What will you use this for?") that routes people to a matching setup. Measure revenue per signup at 90 days and activation.
  3. Templates or sample data vs a blank start. Measure revenue per signup at 90 days and the share who use their own data by day 7.
  4. A setup checklist vs none. Measure revenue per signup and activation.
  5. A product tour vs no tour. Measure time to first value and revenue per signup.
  6. Empty states that show one next action vs a full dashboard. Measure time to first value and revenue per signup.
  7. A day-1 email with a single task vs a general welcome email. Measure activation and revenue per signup.
  8. Asking people to invite teammates in onboarding vs after first value. Measure revenue per account and seats per account.
  9. For apps, the paywall at the end of onboarding vs after the first completed task. See paywall A/B testing.
  10. Onboarding copy that names the first result ("Send your first invoice") vs generic steps ("Complete setup"). Measure time to first value and revenue per signup.

Outtest can test onboarding copy and judges it on revenue from your payment tool rather than on activation alone. Bigger onboarding changes that touch your app's code go through a GitHub pull request you approve. If your onboarding sits inside a free trial, read free trial vs no free trial too, since trial length and onboarding affect each other.

Questions people ask

What is time to first value?+

The time between signing up and the first moment a new user gets something real from the product, such as sending a first invoice or seeing a first report with their own data. Shorter is usually better, but it is a means, not the goal. Measure it as a median, since a few very slow users distort the average.

How do I choose an activation metric?+

Look at past signups and compare how often people paid depending on whether they did each candidate action in their first few days. Pick the action with a big gap in paid rate that enough people reach. Then treat it as a signal to watch, and still judge tests on revenue.

How long should an onboarding test run?+

Long enough to enroll the signups your sample size needs, then long enough for those signups to reach a payment decision, usually the trial length plus one billing cycle. You can read activation within days, but revenue takes 30 to 90 days.

Should onboarding tests be randomized by user or by account?+

By account, if people sign up as teams. If two people in the same workspace see different onboarding, they will confuse each other and the results will blur. For single-user products, randomize by user at signup.

Read next

Let Outtest run your split tests

AI agents read your analytics and payments, find where you lose the most money, build the fix and test it. Every test is judged on revenue, not clicks. Plans from $29 a month.