A/B Testing The Complete Guide
What it actually is, why most tests get called too early, and how to run one that holds up once you switch it on for everyone.
What is A/B testing?
A/B testing is a controlled experiment. You split your traffic in half, send one half to your current page (A, the control) and the other half to a changed version (B, the variant), then compare a specific outcome — sign-ups, purchases, form completions — between the two groups.
It's one method inside the broader discipline of Conversion Rate Optimization: The Complete Guide, which covers research, prioritisation, and the full process of improving what a page does once someone lands on it. A/B testing is the verification step — it's how you find out whether a change you believe will help actually does, instead of guessing.
That distinction matters. Plenty of teams change a page based on a hunch, watch the numbers go up for a week, and call it a win. That's not A/B testing. A/B testing is specifically the practice of testing two versions at the same time, on similarly sized groups of visitors, so external factors — a festival sale, a press mention, a slow Tuesday — hit both versions equally and can't be mistaken for the effect of your change.
Why you can't just eyeball a winner
Because small samples produce misleading swings, and a page you're watching closely will always look like it's doing something in the first few days.
Say your control converts at 3% and your variant is sitting at 4% after two days. That looks like a win. But with a small number of visitors, a handful of extra purchases — pure chance — is enough to create that gap. Flip a coin 20 times and you won't get exactly 10 heads; the same randomness shows up in visitor behaviour.
This is where statistical significance comes in — it's the confidence that a result you're seeing isn't just random noise. A test reaching significance means the gap between A and B is unlikely to have happened by chance alone, given how much data you've collected. A test that hasn't reached significance yet might still turn into a real winner, or might vanish completely as more data comes in — you genuinely can't tell by looking at it.
This is also why sample size matters so much. The fewer visitors you've tested, the wider the range of "normal" random variation — so a gap that looks big early on can shrink to nothing, or reverse, once enough people have seen both versions. More data narrows that range and makes the gap trustworthy. There's no way around needing enough of it.
The three mistakes that quietly invalidate most tests
Almost every broken A/B test fails for one of three reasons, and none of them are visible unless you're specifically watching for them.
| Mistake | What happens | What to do instead |
|---|---|---|
| Stopping the test early because it "looks like" a winner | Early leads are frequently noise. Teams call a winner on day 3–4, roll it out to 100% of traffic, and the lift quietly disappears in production because it was never real to begin with. | Decide your test's stopping point — a significance threshold and a minimum sample — before you launch it, and don't check the result with intent to stop until you get there. |
| Testing trivial things (button colour, icon style) instead of things that move behaviour | You burn weeks of traffic on a change too small to produce a measurable effect either way, and the test ends inconclusive — not because testing failed, but because the change was never big enough to matter. | Test things with real leverage: the offer itself, page structure and content order, form length, pricing presentation, the headline's core promise. |
| Running too many simultaneous tests on the same page or funnel | Visitors overlap between tests, effects interact, and when a metric moves you can't tell which test — or which combination of tests — actually caused it. | Sequence your tests, or if you must run several at once, keep them on pages/funnel stages that don't share traffic, and log what's running where. |
How to run a valid A/B test
The process itself is simple. Getting each step right is what most teams skip.
- Start from a specific problem, not a hunch. Look at where visitors drop off — analytics, session recordings, form abandonment — and pick the one change most likely to affect that drop-off.
- Write down what you expect to happen and why. A one-line hypothesis ("shortening the form will reduce abandonment on mobile") forces clarity before you build anything.
- Decide your success metric and your stopping rules in advance. Pick one primary metric, a significance threshold, and a minimum sample size or run time — and commit to them before the test goes live, not while you're watching the results.
- Split traffic evenly and randomly between A and B. Both versions need to run in the same window, to the same mix of traffic sources and devices, so nothing but the change itself differs between them.
- Let it run to your pre-set stopping point. Resist checking daily with the intent to end it early — checking to monitor is fine, checking to decide isn't, until you've hit your threshold.
- Confirm the result with a proper significance test, not by eye. Run the numbers — visitors and conversions for each version — through a proper test rather than comparing the two percentages directly.
- Document the result either way. A test that shows no difference is still useful information: it tells you that lever isn't the one to pull, and stops the same idea from being retested six months later.
What counts as "enough" sample size and test duration
There's no universal number — it depends on your baseline conversion rate, how much traffic you get, and how big a difference you're actually trying to detect.
A page converting at 8% with heavy traffic can reach a trustworthy result in days. A page converting at 0.5% with modest traffic might need weeks to gather enough conversions to say anything with confidence. The smaller the effect you're hoping to detect, the more data you need to separate it from noise — a change that might lift conversions by a small amount needs a much bigger sample than one you expect to move things a lot.
This is exactly why a rule of thumb ("run it for two weeks", "wait for 1,000 visitors") is the wrong tool. Two different pages with two different baseline rates and two different traffic volumes need two completely different sample sizes to reach the same confidence level. What actually tells you whether a result is real is a proper significance test run on your actual numbers — and you can run the same significance test we use with our free CTR/A-B significance calculator. Put in your visitor and conversion counts for both versions and it tells you whether the gap is statistically real yet, or whether you're still looking at noise.
Multivariate testing: related, but different
Multivariate testing (MVT) tests several elements on the same page at once — say, headline, hero image, and CTA text — in every combination, to see which combination performs best, rather than testing one change at a time like a standard A/B test.
The upside is that MVT can reveal how elements interact with each other, which a sequence of single A/B tests can't show. The downside is traffic: every extra variable multiplies the number of combinations you need to test, which multiplies the traffic required to reach a valid result. MVT only makes sense on high-traffic pages. For most sites, running clean, sequential A/B tests gets you a valid answer far faster than trying to run a multivariate test that never reaches significance.
What actually moves the needle
Structural changes — the offer, the flow, what you're asking a visitor to do and when — consistently give you a bigger, clearer signal to test than surface details like colour or font.
A form that asks for six fields instead of two is a real behavioural lever; a button that's teal instead of blue usually isn't. If you're deciding what to test first, our guide on Multi-Step Forms: When They Help (and When They Hurt) covers one of the highest-leverage structural changes on lead-gen pages, and How to Increase Your Conversion Rate: A Practical Framework walks through how to prioritise which page and which change to test first.
This is also where most in-house testing programmes run out of steam — not from lack of ideas, but from running too many small tests too fast without the discipline to let each one reach a real answer. Our conversion rate optimisation service is built around the opposite approach: fewer tests, run properly, each one backed by research and a defined stopping point, so the result you act on is one you can actually trust.
Frequently asked questions
How long should an A/B test run?
Long enough to reach your pre-set sample size and significance threshold — not a fixed number of days. A high-traffic, high-conversion page might get there in under a week; a low-traffic or low-conversion page can genuinely need several weeks. Set the stopping rule before you launch, based on your actual traffic and baseline rate, rather than picking a duration first and hoping the numbers cooperate by then.
What's a statistically significant result?
It's a result unlikely to have happened by chance alone, given the amount of data collected. In practice, most teams use a 95% confidence level, meaning there's roughly a 5% chance the observed difference is a fluke rather than a genuine effect. A gap that hasn't reached that threshold yet isn't a loss — it just isn't a decision yet either. Running your numbers through a proper significance calculator, like our CTR/A-B significance calculator, is how you find out rather than guessing from the raw percentages.
Can I test more than one thing at once?
You can, but be deliberate about it. Running separate, unrelated A/B tests on pages that don't share traffic is fine. Testing several elements on the same page simultaneously is multivariate testing, and it needs meaningfully more traffic to produce a valid result because it's testing every combination of those elements, not just two versions. On lower-traffic sites, sequential A/B tests almost always reach a trustworthy answer faster than an MVT that never collects enough data per combination.
What if I don't have enough traffic to test?
Then A/B testing on that page probably isn't the right tool yet. Low-traffic pages take a long time to reach a significant sample, and a test that never resolves is wasted effort either way. In that situation, it's usually more productive to make evidence-backed changes — based on user research, session recordings, and structural best practice, like a shorter form or a clearer offer — without running a formal split test, and to save formal A/B testing for the pages that get enough visitors to reach a real answer in a reasonable time.
What's the difference between A/B testing and CRO?
A/B testing is one method used inside CRO. CRO — conversion rate optimisation — is the broader discipline: research, prioritisation, page structure, copy, offers, and testing, all aimed at getting more of your existing traffic to convert. A/B testing is specifically how you validate whether a proposed change works. See What is CRO? Conversion Rate Optimization Explained for the fuller picture of where testing fits into the rest of the process.
Fewer tests. Run properly.
We scope every test against real baseline data and a defined stopping point, so the result you act on is one you can trust.
