A/B Test Sample Size Calculator
Most A/B tests that 'fail' were never capable of succeeding: they were planned with too few visitors to detect the lift they were hunting. Sample size is what fixes that, and the arithmetic is the standard two-proportion formula — sample size grows steeply as the effect you want to detect shrinks relative to the baseline. This calculator applies it with the settings the testing literature treats as defaults: 95% significance, 80% power, two-sided, per variant. It also shows, side by side, what happens when the minimum detectable effect is misread — the single most common planning error in A/B testing, where a relative 10% lift on a 5% baseline is quietly treated as a 10-point jump and the test is planned an order of magnitude (or two) too small. State the effect you would genuinely act on, get the visitor count, divide by your traffic, and you have a test that can actually conclude.
The short answer
Detecting a 10% relative lift on a 5% baseline (5% → 5.5%) at the standard 95% significance and 80% power needs about 31,200 visitors per variant. Read that 10% as absolute points (5% → 15%) and the same formula plans roughly 140 — a 200-fold error that ends the test underpowered with no conclusion.
The control's current conversion rate.
The smallest lift you need the test to catch.
Visitors per variant
31,231
Control and challenger each — 62,462 total. 5.0% → 5.5% is 0.5 points of absolute lift.
- Runtime estimate
- ≈ total ÷ daily traffic
- If 10% were read as absolute
- 138 per variant
Divide 62,462 by the combined visitors both variants get per day — never stop early on a "winning" peek.
The wrong reading: 5.0% → 15% assumes a 10-point jump. Planning this number ends the test underpowered.
How to use A/B Test Sample Size Calculator
- 1
Enter the baseline conversion rate
The control's current rate, measured — not hoped. A 5% baseline is the classic e-commerce figure; landing pages often run 2-3%.
- 2
Set the minimum detectable effect
The smallest lift you would actually act on, marked relative (10% of 5% = 5.5%) or absolute (10 points = 15%). Smaller effects need dramatically more traffic — this is the lever that decides the experiment's length.
- 3
Read the visitors per variant
Each arm needs that many visitors — the total is double. Divide by the combined daily traffic of both variants to get the runtime, and resist peeking at the result until the count is in.
Why use this tool
- Per-variant and total visitor counts from the standard two-proportion formula
- Defaults stated explicitly: 95% significance, 80% power, two-sided
- Relative vs absolute MDE switch — with the misread contrast shown live
- Significance and power adjustable (90/95/99%, 80/90/95% power)
- Runtime reading: divide the total by your daily traffic, not a fixed 'two weeks'
- Free and private — rates never leave your browser
Frequently asked questions
- How many visitors do I need for an A/B test?
- It depends on your baseline and the effect you need to detect. Detecting a 10% relative lift on a 5% baseline (5% → 5.5%) at 95% significance and 80% power needs about 31,200 visitors per variant — 62,400 total. Halve the effect you want to detect and the required sample roughly quadruples.
- What do significance and power actually mean here?
- Significance (95%) is the probability of not declaring a winner when the true conversion rates are identical — the false-positive control. Power (80%) is the probability of detecting the lift if it genuinely exists — the false-negative control. The literature defaults are 95%/80% two-sided; this calculator states them explicitly instead of burying them.
- What is the difference between relative and absolute MDE?
- On a 5% baseline, a 10% relative lift means 5% → 5.5%; a 10-percentage-point absolute lift means 5% → 15%. Reading one for the other is the most common planning error in A/B testing: the absolute misread plans about 140 visitors per variant where the relative reading needs about 31,200 — a test that ends with 'no significant result' purely because it was underpowered.
- Why does a smaller effect need so much more traffic?
- Sample size divides by the square of the effect: detecting 0.5 points takes four times the traffic of detecting 1 point. This is also why high-traffic pages can test subtle UX changes while a newsletter with 2,000 subscribers can only test changes that move conversion by whole percentages.
- How long should my A/B test run?
- Long enough for both variants to reach the required visitor count — total sample divided by combined daily traffic — and in whole business cycles (usually full weeks) so day-of-week effects even out. Run at least two full weeks regardless; never stop early because the peek looks good, which is how false winners are manufactured.
- Does this formula work for more than two variants?
- The visitor count is per pair: a three-way test still needs each challenger compared against the control at full size, so total traffic scales with the number of variants. For multiple comparisons, some testers raise the significance bar (a Bonferroni-style correction) rather than the sample — the direction of both changes is 'more' either way.
- Is my data sent anywhere?
- No. The calculation runs in your browser with the standard published formula — nothing about your baseline, traffic or test is transmitted or stored.
Plan a Google Ads budget with Google's real rules: 30.4× monthly cap, 2× daily overdelivery, clicks, leads, sales and the Smart Bidding floor. Free.
Calculate CPM from spend and impressions, see the viewable CPM behind it, and plan what any budget buys. Free, no sign-up, runs in your browser.
Customer lifetime value done right: gross-margin-adjusted LTV with the revenue figure shown beside it, plus LTV:CAC ratio and CAC payback months.