A/B Test Calculator
Two calculators in one. Before the test: how many visitors each variation needs — and how many whole weeks that is at your traffic, including when the honest answer is that the test is not worth running. After the test: the observed lift, the p-value, a confidence interval and a verdict in plain English.
What you are testing
What the test actually did
How to use the two calculators
- Start on "Before the test" Put in the conversion rate you have now and the smallest lift you would actually act on. Be careful with the relative-versus-absolute switch — this is the single most common error in A/B test planning. "15% better" on a 3% rate means 3.45%. "15 percentage points better" means 18%, which is a different universe and needs about a hundredth of the traffic.
- Read the weeks, not the visitors A number like "24,193 per variation" means nothing on its own. Put your weekly traffic in and the calculator converts it into whole weeks, rounded up, minimum two. That is the number to take into the meeting.
- If the weeks are absurd, use the reverse answer Set the longest you would genuinely run the test and the calculator works backwards: here is the smallest lift you could detect in that window. If it says 32%, you now know that only a substantial redesign is worth testing — not a button colour.
- Switch to "After the test" when it finishes Visitors and conversions for both arms. You get the two rates, the lift, the p-value, the confidence level and — the part that matters most — a confidence interval on the difference, which tells you how big or small the real effect plausibly is.
- Answer the peeking question honestly If you have been checking daily, say so. The tool shows you the upper bound on your real false-positive rate, and it is nowhere near the 5% on the label.
What each input actually controls
| Input | What it means | Effect on sample size |
|---|---|---|
| Baseline rate | The control's true conversion rate | Lower rate, more traffic needed. A 1% baseline needs roughly three times the traffic of a 3% baseline for the same relative lift. |
| Minimum detectable effect | The smallest improvement the test is powered to find | Brutal. Halving the MDE roughly quadruples the sample size — the effect sits in a squared denominator. |
| Significance level (α) | How often you accept a false winner | Tighter is costlier. Moving from 95% to 99% confidence adds roughly half again to the sample — on the worked example below, 24,193 per variation becomes 35,999. |
| Power | How often you catch a real winner | 80% is the convention, and it means you miss one real effect in five. 90% power costs roughly a third more traffic. |
| One or two-sided | Whether a drop counts as a result | One-sided is cheaper and almost always wrong for CRO — you do want to know if the new version is worse. |
The formulae, so you can check the arithmetic
This is a frequentist calculator built on the normal approximation to the binomial. Everything it does is below. Nothing is hidden and there is no server involved — the whole thing runs in your browser.
Sample size per variation — two-proportion test, equal split, pooled variance under the null:
( z(1 − α/s) · √(2 · p̄ · (1 − p̄)) + z(power) · √(p₁(1 − p₁) + p₂(1 − p₂)) )²
n = ───────────────────────────────────────────────────────────────────────────────────
(p₂ − p₁)²
p₁ = baseline conversion rate p₂ = p₁ × (1 + relative MDE) or p₁ + absolute MDE
p̄ = (p₁ + p₂) / 2 s = 2 two-sided, 1 one-sided
z(·) = inverse standard normal CDF n = rounded UP, per variation
total visitors = 2 × n
Significance after the test — pooled two-proportion z-test:
p₁ = c₁ / n₁ p₂ = c₂ / n₂ p_pool = (c₁ + c₂) / (n₁ + n₂)
SE_pool = √( p_pool · (1 − p_pool) · (1/n₁ + 1/n₂) )
z = (p₂ − p₁) / SE_pool
p-value = 2 · (1 − Φ(|z|)) two-sided
= 1 − Φ(z) one-sided, in the declared direction
confidence = (1 − p-value) × 100%
Confidence interval on the difference — Wald interval, unpooled standard error. The test above pools the two proportions because it is testing the hypothesis that they are equal; the interval does not pool, because it is estimating a difference that may well be real. That mismatch is standard practice and is worth knowing about rather than discovering later:
SE_diff = √( p₁(1 − p₁)/n₁ + p₂(1 − p₂)/n₂ ) CI = (p₂ − p₁) ± z(1 − α/2) · SE_diff relative CI = CI bounds ÷ p₁ (the lift range, as a percentage of baseline)
Duration and the peeking bound:
weeks = ceil( (2 × n) ÷ weekly visitors entering the test ), minimum 2
peeking, k independent looks at level α: 1 − (1 − α)^k
This is an UPPER BOUND, not the real number. Successive looks at the same
accumulating test are heavily correlated, so the true inflated false-positive
rate sits somewhere between α and this bound — but it is never α.
The inverse normal CDF is Acklam's rational approximation followed by one Halley refinement step against a Hart/West normal CDF, which is accurate to roughly one part in 1015 — it returns 1.9599639845400538 for the 97.5th percentile against a true value of 1.9599639845400545. The arithmetic is not where this page will let you down.
Two worked examples you can reproduce
Both of these are the defaults in the tool, so you can check them without typing anything.
| Example | Inputs | Result |
|---|---|---|
| Sample size | Baseline 3.00%, +15% relative (target 3.45%), α = 0.05 two-sided, power 80%, 8,000 visitors/week | 24,193 per variation · 48,386 total · 7 whole weeks |
| Significance | Control 8,000 visitors / 240 conversions (3.00%), variation 8,000 / 300 (3.75%), α = 0.05 two-sided | +25.00% relative lift · z = 2.6267 · p = 0.0086 · 95% CI on the difference +0.19pp to +1.31pp, i.e. a +6.35% to +43.65% lift |
Look hard at that second interval. The test is significant, the point estimate is a 25% lift, and the data is equally consistent with a 6% lift as with a 44% one. If your business case only works at 20%, a significant result is not the same as a green light. This is why the interval is on the page and not buried behind an "advanced" toggle.
Why whole weeks, and why peeking breaks the test
Traffic has a weekly rhythm. B2B enquiries collapse at weekends, D2C carts spike on Sunday nights, and in India the first week after salaries land does not look like the last week of the month. A test that runs 11 days contains two Mondays and one Saturday, so it has silently over-weighted one kind of visitor. Running in whole weeks — and at least two of them, so you have seen each day twice — removes that distortion without any statistics. It is the cheapest correctness win in testing.
Peeking is the more expensive mistake. A p-value at the 95% threshold is a promise about one pre-declared comparison: if there were truly no difference, you would see a result this extreme 5% of the time. Check the test every morning and stop the first time it crosses, and you have not run one comparison — you have run as many as you have looked, and kept only the most flattering one. With k looks the chance of at least one false crossing rises towards 1 − (1 − α)k; the looks are correlated so the real figure is lower than that bound, but it is always above the α you think you are working at. The fix is not statistical sophistication. It is deciding the sample size first, writing it down, and not touching the report until it is reached.
What this calculator cannot do
Four real limits, and the last one is the one that costs money.
- It is frequentist and approximate. Normal approximation to the binomial — no Bayesian posteriors, no "probability B beats A", no sequential or always-valid testing, no multi-armed bandits. If your testing tool reports a "chance to win" percentage, that is a Bayesian output and it will not match the p-value here, because it is answering a different question.
- It breaks at low conversion counts. The approximation needs at least 10 conversions and 10 non-conversions in each arm as a hard floor, and in practice you want 25 or more conversions per arm before the numbers deserve a decision. The tool flags this rather than silently returning a confident-looking p-value on nine conversions. Nine conversions against six is noise, whatever the maths says.
- It handles one variation, on a binary metric. Two arms, one conversion event. Test three variations against a control and you have three comparisons, which needs a multiple-comparison correction this tool does not apply — run them one at a time or tighten α yourself. Revenue per visitor, average order value and time on page are continuous metrics and need a different test entirely.
- It cannot see your test. It does not know whether your traffic split was actually even, whether the same visitor got both variations across devices, whether a deploy went out mid-test, or whether the conversion event fires correctly on both pages. Every one of those has produced more false winners than bad maths ever has. If the two arms received wildly different visitor counts when you meant to split 50/50, stop and fix the instrumentation before you trust any number on this page.
One more thing worth saying plainly: a conversion-rate win is not automatically a revenue win. A variation can lift signups and attract worse-quality ones. Check the lift against what those customers are worth — our note on customer lifetime value covers the numbers to look at before you roll a winner out.
Where to go next
This page does the arithmetic. What to test, how to build a hypothesis worth testing and how to run a programme rather than a series of one-off experiments is the A/B testing guide, which sits inside the conversion rate optimisation guide. If you just need the rate itself rather than a test, the conversion rate calculator is next door, and the rest of the free tools cover the neighbouring jobs. When you want someone to run the testing programme with you, see how we work.
Frequently asked questions
How do I calculate the sample size for an A/B test?
You need four numbers: your current conversion rate, the smallest improvement worth detecting, your significance level (usually 95%) and your statistical power (usually 80%). The sample size per variation is then the squared sum of two z-scaled standard deviations divided by the squared difference between the two rates — the exact formula is printed on this page. The important part is what it implies: because the difference sits in a squared denominator, halving the effect you want to detect roughly quadruples the traffic you need.
What is the difference between relative and absolute minimum detectable effect?
This is the error that ruins more test plans than any other. On a 3% baseline, a 15% relative improvement is 3.45% — a 0.45 percentage point move. A 15 percentage point absolute improvement is 18%, which is a five-fold increase. The second needs a tiny fraction of the traffic because the effect is enormous. If a calculator gives you a suspiciously small sample size, check which one it assumed. This one makes you choose.
My result is not statistically significant. Does that mean the two versions are the same?
No, and this is the most consequential misreading in testing. "Not significant" means the test did not gather enough evidence to rule out chance — it says nothing about the size of the real effect. Read the confidence interval instead. If it runs from a 2% drop to a 3% lift, you genuinely can act as though the difference is small. If it runs from a 12% drop to a 20% lift, your test simply did not have the traffic to answer the question, and calling that a tie is inventing information you do not have.
Can I stop the test early if it already shows 95% confidence?
Not without paying for it. The 5% false-positive rate a 95% threshold promises applies to a single pre-declared check. Every additional look is another chance for noise to cross the line, and stopping at the first crossing keeps only the most flattering moment of the test. The upper bound on your true false-positive rate after k looks is 1 − (1 − α)k; the real figure is lower than that because the looks are correlated, but it is never the 5% on the label. Fix the sample size in advance, then look once. If you genuinely need to stop early, that requires sequential testing methods, which this calculator does not implement.
Is anything I type in here sent to you?
No. Every calculation happens in your browser as you type. Nothing is uploaded, nothing is stored and nothing is remembered between visits — refresh the page and it is gone. The tool works with the network disconnected once the page has loaded.
Most sites do not need a test. They need the obvious thing fixed first.
Send us your site and your traffic. We will tell you what is worth testing, what is too small to ever prove, and what to just go and change.
