Home › Free Tools › A/B Test Calculator
FREE TOOL · NO SIGNUP

A/B Test Calculator

Two calculators in one. Before the test: how many visitors each variation needs — and how many whole weeks that is at your traffic, including when the honest answer is that the test is not worth running. After the test: the observed lift, the p-value, a confidence interval and a verdict in plain English.

WeeksNot just a visitor count 0Data leaves your browser FreeNo signup, no limits
By the Digital Hangover team · Updated September 2026 · Free forever

What you are testing

The control's real rate. Take it from at least a month of data, not a good week.
The smallest improvement worth finding. Relative means "15% better than 3%, so 3.45%". Absolute means "3% becomes 4%".
The false-positive rate you accept.
Your chance of spotting a real effect.
Two-sided unless you genuinely do not care about a drop.
Total across both variations — not per variation, and not your whole site if only some pages are in the test.
Used for the reverse answer: the smallest lift you could actually detect in that window.
Quick answer: an A/B test needs a sample size fixed before it starts, and that sample size is driven almost entirely by two numbers — your current conversion rate and the smallest improvement you would act on. Small lifts on low rates need enormous traffic. At a 3% conversion rate, finding a 15% relative improvement takes roughly 24,000 visitors per variation at the usual 95% / 80% settings. At 8,000 visitors a week that is seven whole weeks. If that sentence just killed your test, that is the calculator doing its job.

How to use the two calculators

  1. Start on "Before the test" Put in the conversion rate you have now and the smallest lift you would actually act on. Be careful with the relative-versus-absolute switch — this is the single most common error in A/B test planning. "15% better" on a 3% rate means 3.45%. "15 percentage points better" means 18%, which is a different universe and needs about a hundredth of the traffic.
  2. Read the weeks, not the visitors A number like "24,193 per variation" means nothing on its own. Put your weekly traffic in and the calculator converts it into whole weeks, rounded up, minimum two. That is the number to take into the meeting.
  3. If the weeks are absurd, use the reverse answer Set the longest you would genuinely run the test and the calculator works backwards: here is the smallest lift you could detect in that window. If it says 32%, you now know that only a substantial redesign is worth testing — not a button colour.
  4. Switch to "After the test" when it finishes Visitors and conversions for both arms. You get the two rates, the lift, the p-value, the confidence level and — the part that matters most — a confidence interval on the difference, which tells you how big or small the real effect plausibly is.
  5. Answer the peeking question honestly If you have been checking daily, say so. The tool shows you the upper bound on your real false-positive rate, and it is nowhere near the 5% on the label.

What each input actually controls

InputWhat it meansEffect on sample size
Baseline rateThe control's true conversion rateLower rate, more traffic needed. A 1% baseline needs roughly three times the traffic of a 3% baseline for the same relative lift.
Minimum detectable effectThe smallest improvement the test is powered to findBrutal. Halving the MDE roughly quadruples the sample size — the effect sits in a squared denominator.
Significance level (α)How often you accept a false winnerTighter is costlier. Moving from 95% to 99% confidence adds roughly half again to the sample — on the worked example below, 24,193 per variation becomes 35,999.
PowerHow often you catch a real winner80% is the convention, and it means you miss one real effect in five. 90% power costs roughly a third more traffic.
One or two-sidedWhether a drop counts as a resultOne-sided is cheaper and almost always wrong for CRO — you do want to know if the new version is worse.

The formulae, so you can check the arithmetic

This is a frequentist calculator built on the normal approximation to the binomial. Everything it does is below. Nothing is hidden and there is no server involved — the whole thing runs in your browser.

Sample size per variation — two-proportion test, equal split, pooled variance under the null:

        ( z(1 − α/s) · √(2 · p̄ · (1 − p̄))  +  z(power) · √(p₁(1 − p₁) + p₂(1 − p₂)) )²
n  =   ───────────────────────────────────────────────────────────────────────────────────
                                        (p₂ − p₁)²

  p₁ = baseline conversion rate            p₂ = p₁ × (1 + relative MDE)   or   p₁ + absolute MDE
  p̄  = (p₁ + p₂) / 2                       s  = 2 two-sided, 1 one-sided
  z(·) = inverse standard normal CDF        n  = rounded UP, per variation
  total visitors = 2 × n

Significance after the test — pooled two-proportion z-test:

  p₁ = c₁ / n₁          p₂ = c₂ / n₂          p_pool = (c₁ + c₂) / (n₁ + n₂)

  SE_pool = √( p_pool · (1 − p_pool) · (1/n₁ + 1/n₂) )

  z = (p₂ − p₁) / SE_pool

  p-value = 2 · (1 − Φ(|z|))     two-sided
          = 1 − Φ(z)             one-sided, in the declared direction
  confidence = (1 − p-value) × 100%

Confidence interval on the difference — Wald interval, unpooled standard error. The test above pools the two proportions because it is testing the hypothesis that they are equal; the interval does not pool, because it is estimating a difference that may well be real. That mismatch is standard practice and is worth knowing about rather than discovering later:

  SE_diff = √( p₁(1 − p₁)/n₁  +  p₂(1 − p₂)/n₂ )

  CI = (p₂ − p₁)  ±  z(1 − α/2) · SE_diff

  relative CI = CI bounds ÷ p₁        (the lift range, as a percentage of baseline)

Duration and the peeking bound:

  weeks = ceil( (2 × n) ÷ weekly visitors entering the test ),  minimum 2

  peeking, k independent looks at level α:   1 − (1 − α)^k

    This is an UPPER BOUND, not the real number. Successive looks at the same
    accumulating test are heavily correlated, so the true inflated false-positive
    rate sits somewhere between α and this bound — but it is never α.

The inverse normal CDF is Acklam's rational approximation followed by one Halley refinement step against a Hart/West normal CDF, which is accurate to roughly one part in 1015 — it returns 1.9599639845400538 for the 97.5th percentile against a true value of 1.9599639845400545. The arithmetic is not where this page will let you down.

Two worked examples you can reproduce

Both of these are the defaults in the tool, so you can check them without typing anything.

ExampleInputsResult
Sample size Baseline 3.00%, +15% relative (target 3.45%), α = 0.05 two-sided, power 80%, 8,000 visitors/week 24,193 per variation · 48,386 total · 7 whole weeks
Significance Control 8,000 visitors / 240 conversions (3.00%), variation 8,000 / 300 (3.75%), α = 0.05 two-sided +25.00% relative lift · z = 2.6267 · p = 0.0086 · 95% CI on the difference +0.19pp to +1.31pp, i.e. a +6.35% to +43.65% lift

Look hard at that second interval. The test is significant, the point estimate is a 25% lift, and the data is equally consistent with a 6% lift as with a 44% one. If your business case only works at 20%, a significant result is not the same as a green light. This is why the interval is on the page and not buried behind an "advanced" toggle.

Why whole weeks, and why peeking breaks the test

Traffic has a weekly rhythm. B2B enquiries collapse at weekends, D2C carts spike on Sunday nights, and in India the first week after salaries land does not look like the last week of the month. A test that runs 11 days contains two Mondays and one Saturday, so it has silently over-weighted one kind of visitor. Running in whole weeks — and at least two of them, so you have seen each day twice — removes that distortion without any statistics. It is the cheapest correctness win in testing.

Peeking is the more expensive mistake. A p-value at the 95% threshold is a promise about one pre-declared comparison: if there were truly no difference, you would see a result this extreme 5% of the time. Check the test every morning and stop the first time it crosses, and you have not run one comparison — you have run as many as you have looked, and kept only the most flattering one. With k looks the chance of at least one false crossing rises towards 1 − (1 − α)k; the looks are correlated so the real figure is lower than that bound, but it is always above the α you think you are working at. The fix is not statistical sophistication. It is deciding the sample size first, writing it down, and not touching the report until it is reached.

What this calculator cannot do

Four real limits, and the last one is the one that costs money.

  • It is frequentist and approximate. Normal approximation to the binomial — no Bayesian posteriors, no "probability B beats A", no sequential or always-valid testing, no multi-armed bandits. If your testing tool reports a "chance to win" percentage, that is a Bayesian output and it will not match the p-value here, because it is answering a different question.
  • It breaks at low conversion counts. The approximation needs at least 10 conversions and 10 non-conversions in each arm as a hard floor, and in practice you want 25 or more conversions per arm before the numbers deserve a decision. The tool flags this rather than silently returning a confident-looking p-value on nine conversions. Nine conversions against six is noise, whatever the maths says.
  • It handles one variation, on a binary metric. Two arms, one conversion event. Test three variations against a control and you have three comparisons, which needs a multiple-comparison correction this tool does not apply — run them one at a time or tighten α yourself. Revenue per visitor, average order value and time on page are continuous metrics and need a different test entirely.
  • It cannot see your test. It does not know whether your traffic split was actually even, whether the same visitor got both variations across devices, whether a deploy went out mid-test, or whether the conversion event fires correctly on both pages. Every one of those has produced more false winners than bad maths ever has. If the two arms received wildly different visitor counts when you meant to split 50/50, stop and fix the instrumentation before you trust any number on this page.

One more thing worth saying plainly: a conversion-rate win is not automatically a revenue win. A variation can lift signups and attract worse-quality ones. Check the lift against what those customers are worth — our note on customer lifetime value covers the numbers to look at before you roll a winner out.

Where to go next

This page does the arithmetic. What to test, how to build a hypothesis worth testing and how to run a programme rather than a series of one-off experiments is the A/B testing guide, which sits inside the conversion rate optimisation guide. If you just need the rate itself rather than a test, the conversion rate calculator is next door, and the rest of the free tools cover the neighbouring jobs. When you want someone to run the testing programme with you, see how we work.

Frequently asked questions

How do I calculate the sample size for an A/B test?

You need four numbers: your current conversion rate, the smallest improvement worth detecting, your significance level (usually 95%) and your statistical power (usually 80%). The sample size per variation is then the squared sum of two z-scaled standard deviations divided by the squared difference between the two rates — the exact formula is printed on this page. The important part is what it implies: because the difference sits in a squared denominator, halving the effect you want to detect roughly quadruples the traffic you need.

What is the difference between relative and absolute minimum detectable effect?

This is the error that ruins more test plans than any other. On a 3% baseline, a 15% relative improvement is 3.45% — a 0.45 percentage point move. A 15 percentage point absolute improvement is 18%, which is a five-fold increase. The second needs a tiny fraction of the traffic because the effect is enormous. If a calculator gives you a suspiciously small sample size, check which one it assumed. This one makes you choose.

My result is not statistically significant. Does that mean the two versions are the same?

No, and this is the most consequential misreading in testing. "Not significant" means the test did not gather enough evidence to rule out chance — it says nothing about the size of the real effect. Read the confidence interval instead. If it runs from a 2% drop to a 3% lift, you genuinely can act as though the difference is small. If it runs from a 12% drop to a 20% lift, your test simply did not have the traffic to answer the question, and calling that a tie is inventing information you do not have.

Can I stop the test early if it already shows 95% confidence?

Not without paying for it. The 5% false-positive rate a 95% threshold promises applies to a single pre-declared check. Every additional look is another chance for noise to cross the line, and stopping at the first crossing keeps only the most flattering moment of the test. The upper bound on your true false-positive rate after k looks is 1 − (1 − α)k; the real figure is lower than that because the looks are correlated, but it is never the 5% on the label. Fix the sample size in advance, then look once. If you genuinely need to stop early, that requires sequential testing methods, which this calculator does not implement.

Is anything I type in here sent to you?

No. Every calculation happens in your browser as you type. Nothing is uploaded, nothing is stored and nothing is remembered between visits — refresh the page and it is gone. The tool works with the network disconnected once the page has loaded.

FREE CONVERSION REVIEW

Most sites do not need a test. They need the obvious thing fixed first.

Send us your site and your traffic. We will tell you what is worth testing, what is too small to ever prove, and what to just go and change.

See how we work →