Data & Product Analytics
The 10% Lift That Wasn't: How A/B Tests Fool Product Teams
Your new checkout button goes live. After three days, conversions are up 10% and the dashboard glows green. Do you ship it? Product teams face this moment constantly, and many get it wrong, because noise and impatience quietly corrupt the result.

A lift is not yet evidence
Say the old page converts 4.0% of visitors and the new one 4.4%, with 5,000 visitors each. That is a 10% relative lift. The difference of two proportions calculator returns a p-value near 0.32, fully consistent with chance. The two-proportion confidence interval shows why: the true difference could plausibly range from a small loss to a decent gain.
Keep the same rates but collect 50,000 visitors per group, and the p-value falls to about 0.002.
The effect did not change. The evidence did.
This is the heart of the problem. Dashboards display a single tidy percentage, which hides the uncertainty around it. A 10% lift measured on a few thousand visitors is a noisy estimate, and noisy estimates swing wildly in both directions. Teams then celebrate the lucky swings and quietly forget the unlucky ones.
Decide the sample size before you start
Detecting a move from 4.0% to 4.4% with 80% power at the 5% level takes roughly 39,500 visitors per group. As a rule of thumb, halving the effect you want to detect roughly quadruples the traffic you need. Small improvements are expensive to confirm, and it is better to know that upfront than to discover it after launch.
The peeking problem
Checking the dashboard daily and stopping once p dips below 0.05 inflates false positives dramatically. With five looks, the true error rate is closer to 14% than the advertised 5%. Either fix the sample size and look once, or use methods designed for sequential testing.
There is a second trap hiding in plain sight: testing many variants or many metrics at once. Compare twenty button colours and, at the 5% level, one will usually look like a winner by chance alone. Pre-registering a single main metric before the test starts is a simple, powerful defence.
Check the split before the result
Before reading any outcome, confirm that traffic was divided as intended. Suppose a 50/50 test shows 5,200 visitors in one group and 4,800 in the other. That looks harmless, but the chi-square goodness-of-fit calculator gives a p-value far below 0.001, a sign that something in the assignment or tracking is broken. Teams call this a sample ratio mismatch, and it invalidates the comparison.
Timing matters as well. Novelty can inflate early results because people click on anything new, and weekday habits differ from weekends. Running a test for whole weeks, rather than stopping on a lucky Tuesday, protects against both distortions.
Decide what counts as a win
Before launch, write down the primary metric, the smallest lift worth shipping, and the sample size. A tiny lift can be statistically real yet too small to justify engineering effort or extra complexity. Add guardrail metrics as well, such as refund rate or page speed, because a change that boosts clicks while quietly hurting retention is not a win.
Finally, share the negative results. A test that finds no difference is still information: it saves other teams from repeating the experiment and keeps everyone honest about how often ideas fail. Many do.
Not every metric is a percentage
For revenue per visitor or time on page, compare averages with the independent samples t-test calculator. If the data is heavily skewed, as spending usually is, the Mann-Whitney test is a more forgiving choice.