ARTICLE

Why most A/B tests fail before they even start

CRO

Run the numbers on your own testing program and the picture is usually grim. In an analysis of 28,304 experiments published by CXL, only one in five reached 95% statistical significance. Optimizely reported a similar rate across 127,000 experiments, around 12%. The rest were inconclusive, stopped too early, or too weak to prove anything.

The instinct is to blame the test: wrong button, wrong headline. The real problem sits earlier, in the decisions made before anyone clicks "start".

The first is traffic. Statistical significance is not a switch you flip; it needs volume. A rough working figure is a few thousand visitors and around a hundred conversions per variation before a result means anything. Below that, you are not running an experiment, you are collecting anecdotes. Linear’s 2026 review is blunt about it: under 1,000 monthly conversions, the signal is too noisy to trust. Most teams testing on low-traffic pages have already lost before they begin. The fix is not a better tool. It is testing on the pages that actually get traffic, and accepting that some pages will never have enough.

The second is the hypothesis. A test is only as good as the reason behind it. "Let’s try a green button" is not a hypothesis, it is a guess wearing a lab coat. A real hypothesis comes from somewhere: a drop-off you saw in the funnel, a pattern in session replays, a complaint that keeps recurring. The programs that pre-qualify ideas with analytics, heatmaps and user research win far more often than those that test whatever lands in a Slack thread. The research is not a nice-to-have before the test. It is the test’s foundation.

Then there is the mistake everyone makes at least once: stopping early. Your tool shows 95% significance on day three and a 25% lift, so you call it. Here is the uncomfortable fact. Roughly three quarters of A/A tests, where you pit a page against an identical copy of itself, will cross 95% significance at some point if you keep looking. Significance reached by peeking is noise dressed as a result. CXL calls stopping a test the moment it looks significant the deadliest habit in the discipline, and they are right. The lift you "won" evaporates on rollout, and worse, you now believe something false and carry it into the next decision.

None of this is about running fewer tests. It is about respecting what a test is. Before you launch, three questions decide the outcome more than the variant ever will. Does this page have the traffic to reach significance in a reasonable window? Is there a real observation behind this idea, or am I guessing? And have I fixed a sample size and duration in advance, so I am not tempted to stop the moment the graph flatters me?

Answer those honestly and a lot of tests never launch. That is the point. The ones that do are the ones worth your time. The teams that win at experimentation are not the ones running the most tests. They are the ones who did the thinking before the test existed.

43

Share It

El Mahdi Khiyat

Digital analyst who reads the data, shapes the experience, and builds the page that answers it.

PARIS · GMT+1

© 2026 El Mahdi Khiyat. All rights reserved.

I care about the person behind the click.