Experiments & Paired Data
Did the New Habit Really Work? Making Sense of Before-and-After Data
You start walking after dinner for a month and your step count rises. You adopt a new sleep routine and feel sharper. Before-and-after comparisons are the most natural way to judge a change, and one of the easiest to analyze incorrectly.

Why pairing matters
Imagine eight people tracking daily steps, in thousands, before and after a challenge. Their average rises from 6.7 to 7.3, but everyone is different: some walk a lot to begin with, others little. The comparison that matters is within each person, so each individual's change is the real data point.
Pairing protects against another trap. Comparing the average "before" with the average "after" can bury a consistent small gain under large differences between people. The within-person difference is a much cleaner signal, and it needs fewer participants to detect an effect of the same size.
Treating the two sets as unrelated groups ignores that structure. The independent samples t-test calculator on these numbers gives a p-value around 0.12 and calls the effect unconvincing. The paired samples t-test calculator, which analyzes each person's difference (mean 0.64, standard deviation 0.33), gives about 0.001.
Same data, opposite conclusions, because pairing removes the variation between people.
When the data is not normal
With only eight people, you cannot easily verify that the differences follow a bell curve. The Wilcoxon signed-rank calculator for paired data uses only the ranks of the changes. Here all eight changes are positive, which gives an exact two-sided p-value of about 0.008, still strong evidence of a shift.
Add an interval for the size of the change. Feeding the eight differences into the confidence interval calculator for unknown variance gives an average gain of roughly 0.36 to 0.91 thousand steps a day, a range that may be modest in practical terms even though it is statistically clear.
What the numbers cannot tell you
A significant paired result shows that scores changed, not why. Without a control group, other explanations remain: seasons change, people join challenges when motivated, and being observed can nudge behavior. Extreme starting points also tend to drift toward the average on their own, a phenomenon called regression to the mean.
Stronger designs add a comparison group that does not change its habit, or randomize who starts first. Even a simple wait-list group helps separate the effect of the habit from the effect of simply taking part.
Practical tips for tidy data
Collect data the same way before and after, ideally on the same device and at a similar time of year, since a new watch can change readings by itself. Record a baseline long enough to be representative: one week may be unusual, while two or three weeks give a steadier reference. Keep a short diary of other changes, such as travel or illness, that might affect the numbers.
Report the p-value and the interval, and plot a line from each person's before value to their after value. If most lines slope upward, readers see the story at a glance. If they scatter, the average is hiding inconsistency.
Remember that the mean gain of 0.64 hides individual differences: one person gained only 0.1 while another gained 1.1. Eight people suffice for a demonstration, not for a strong claim. Decide your main outcome before collecting data, rather than highlighting whichever of steps, sleep or mood happened to look significant.
Before celebrating a "before and after" chart, ask four questions: are the data paired, how large is the typical change, how uncertain is that estimate, and what else could explain it?