
Why Most A/B Tests Conclude Before They Reach Significance
If your traffic is below 10K visits per variant, you are reading noise. The fix.
Most A/B tests are called too early and the winner is noise. The decision then gets implemented, believed, and quoted for years.
The mechanism is simple. Early in a test the numbers swing wildly, and if you watch them daily you will always see a moment where one variant looks clearly ahead.
Why early leads are meaningless
With a small number of conversions, one extra sale moves the conversion rate by a large amount. That volatility is entirely random, and it settles as volume grows.
Stopping the moment a variant looks good is not a shortcut. It is selecting for the random peak, which is why so many implemented winners fail to reproduce their lift afterwards.
The rough scale of what you need
The lower your conversion rate and the smaller the effect you are trying to detect, the more traffic it takes.
A small improvement on a low-converting page needs a genuinely large sample per variant. A large improvement on a page that converts well can be detected with far less. Detecting a two percent relative lift is not a realistic goal for most sites, and chasing it wastes months.
Run the numbers before the test rather than after, using any sample size calculator. If the required sample exceeds a month or two of traffic, the test is not viable at that effect size and you should be testing something bigger.
Fix the duration before you start
Two full weeks minimum, and always in whole weeks.
Weekday and weekend traffic behave differently, and a test ending on a Wednesday has an unbalanced sample. Whole weeks remove that.
Two weeks also covers the difference between people who convert immediately and people who return days later.
Decide the rules in advance
Write down the sample size, the duration, the primary metric and what result changes what action. Before the test runs.
The alternative is deciding those things while looking at the results, which is how a null result becomes a story about a segment that happened to look good.
One primary metric. Secondary metrics are for understanding, not for declaring a winner.
What to do when you do not have the traffic
Most sites do not have enough traffic for meaningful split testing, and pretending otherwise wastes the year.
Test large changes rather than small ones. A different offer, a different page structure, a different core message. Big effects are detectable with small samples, and the marginal ones you cannot measure are also the ones that would not have mattered.
Fix the known problems instead. Slow load, a form asking for eleven fields, no proof near the call to action, a headline that does not say what you do. Those do not need a test, they need doing.
Use qualitative evidence. Session recordings, exit surveys, five user tests. Ten people struggling with the same step tells you more than an underpowered test on the same page.
Test sequentially, not simultaneously
Running several overlapping tests on the same funnel makes each result unreadable, because you cannot separate the effects.
One test at a time on any given path, run to its planned size, then the next.
The honest summary
If your traffic is low, split testing is not your growth lever. Fixing the obvious and shipping bigger changes is. Testing is a method for refining something that already works, and it needs volume to be worth anything at all.
Written by David Eid. Published .
Read next.
Contact the Ignis Team
Send through your details and we will audit your business before we reply.




