ToolNest

A/B Test Sample Size Calculator

Most A/B tests that 'fail' were never capable of succeeding: they were planned with too few visitors to detect the lift they were hunting. Sample size is what fixes that, and the arithmetic is the standard two-proportion formula — sample size grows steeply as the effect you want to detect shrinks relative to the baseline. This calculator applies it with the settings the testing literature treats as defaults: 95% significance, 80% power, two-sided, per variant. It also shows, side by side, what happens when the minimum detectable effect is misread — the single most common planning error in A/B testing, where a relative 10% lift on a 5% baseline is quietly treated as a 10-point jump and the test is planned an order of magnitude (or two) too small. State the effect you would genuinely act on, get the visitor count, divide by your traffic, and you have a test that can actually conclude.

The short answer

Detecting a 10% relative lift on a 5% baseline (5% → 5.5%) at the standard 95% significance and 80% power needs about 31,200 visitors per variant. Read that 10% as absolute points (5% → 15%) and the same formula plans roughly 140 — a 200-fold error that ends the test underpowered with no conclusion.

The control's current conversion rate.

The smallest lift you need the test to catch.

Visitors per variant

31,231

Control and challenger each — 62,462 total. 5.0% → 5.5% is 0.5 points of absolute lift.

Runtime estimate
≈ total ÷ daily traffic

Divide 62,462 by the combined visitors both variants get per day — never stop early on a "winning" peek.

If 10% were read as absolute
138 per variant

The wrong reading: 5.0% → 15% assumes a 10-point jump. Planning this number ends the test underpowered.

How to use A/B Test Sample Size Calculator

  1. 1

    Enter the baseline conversion rate

    The control's current rate, measured — not hoped. A 5% baseline is the classic e-commerce figure; landing pages often run 2-3%.

  2. 2

    Set the minimum detectable effect

    The smallest lift you would actually act on, marked relative (10% of 5% = 5.5%) or absolute (10 points = 15%). Smaller effects need dramatically more traffic — this is the lever that decides the experiment's length.

  3. 3

    Read the visitors per variant

    Each arm needs that many visitors — the total is double. Divide by the combined daily traffic of both variants to get the runtime, and resist peeking at the result until the count is in.

Why use this tool

  • Per-variant and total visitor counts from the standard two-proportion formula
  • Defaults stated explicitly: 95% significance, 80% power, two-sided
  • Relative vs absolute MDE switch — with the misread contrast shown live
  • Significance and power adjustable (90/95/99%, 80/90/95% power)
  • Runtime reading: divide the total by your daily traffic, not a fixed 'two weeks'
  • Free and private — rates never leave your browser

Frequently asked questions

How many visitors do I need for an A/B test?
It depends on your baseline and the effect you need to detect. Detecting a 10% relative lift on a 5% baseline (5% → 5.5%) at 95% significance and 80% power needs about 31,200 visitors per variant — 62,400 total. Halve the effect you want to detect and the required sample roughly quadruples.
What do significance and power actually mean here?
Significance (95%) is the probability of not declaring a winner when the true conversion rates are identical — the false-positive control. Power (80%) is the probability of detecting the lift if it genuinely exists — the false-negative control. The literature defaults are 95%/80% two-sided; this calculator states them explicitly instead of burying them.
What is the difference between relative and absolute MDE?
On a 5% baseline, a 10% relative lift means 5% → 5.5%; a 10-percentage-point absolute lift means 5% → 15%. Reading one for the other is the most common planning error in A/B testing: the absolute misread plans about 140 visitors per variant where the relative reading needs about 31,200 — a test that ends with 'no significant result' purely because it was underpowered.
Why does a smaller effect need so much more traffic?
Sample size divides by the square of the effect: detecting 0.5 points takes four times the traffic of detecting 1 point. This is also why high-traffic pages can test subtle UX changes while a newsletter with 2,000 subscribers can only test changes that move conversion by whole percentages.
How long should my A/B test run?
Long enough for both variants to reach the required visitor count — total sample divided by combined daily traffic — and in whole business cycles (usually full weeks) so day-of-week effects even out. Run at least two full weeks regardless; never stop early because the peek looks good, which is how false winners are manufactured.
Does this formula work for more than two variants?
The visitor count is per pair: a three-way test still needs each challenger compared against the control at full size, so total traffic scales with the number of variants. For multiple comparisons, some testers raise the significance bar (a Bonferroni-style correction) rather than the sample — the direction of both changes is 'more' either way.
Is my data sent anywhere?
No. The calculation runs in your browser with the standard published formula — nothing about your baseline, traffic or test is transmitted or stored.

Related tools