Research Methods
Coffee, Screens and Sleep: Why a Strong Correlation Still Doesn't Prove Cause
"Study links late-night scrolling to lower grades." Headlines like this appear weekly, and nearly all of them rest on one number: the correlation coefficient. It is a valuable number, and one of the easiest to misread.

What r actually tells you
Pearson's r runs from −1 to +1 and measures how closely points follow a straight line. Square it and you get the share of variation the two variables have in common: an r of 0.8 means about 64%. Try your own data in the Pearson correlation calculator, or build intuition with the correlation coefficient simulator, where moving a single point shows how an outlier can swing the result.
Same r, completely different pictures
In 1973, statistician Francis Anscombe published four small datasets with nearly identical means, variances and correlations of about 0.82. Plotted, they look nothing alike: one is a clean line, one a curve, and in one a single extreme point creates the whole relationship. Always plot first.
When data is ranked, skewed, or contains outliers, rank-based measures are sturdier. The Spearman correlation calculator and the Kendall tau and Spearman calculator both work on ranks rather than raw values.
The missing third variable
Ice cream sales and drownings rise together every summer, yet nobody blames ice cream. Hot weather drives both. Real examples are subtler. Students who sleep less may also work longer hours, carry more stress, or live farther from campus, and any of these could explain a link with grades.
No correlation, however significant, can separate those explanations. A simple linear regression shows how well one variable predicts another, and multivariate regression lets you add the other factors you suspect. But prediction is still not causation. Only randomized experiments, or carefully designed studies that account for confounders, move us closer to cause.
Direction can also be backwards. Do screens harm sleep, or do restless sleepers reach for their phones at 2 a.m.? A correlation is silent about which one leads, and both may be partly true.
Reverse causation is one of the most common blind spots in popular science reporting.
Small samples, big coincidences
Sample size changes what a number means. With only 10 people, an r near 0.63 is needed before the result clears the usual 5% significance bar, so impressive-looking correlations appear by luck all the time. With 100 people, an r of just 0.20 is already significant, yet it explains only 4% of the variation. Significant and meaningful are different things.
Then there is data dredging. Test enough pairs of variables and roughly one in twenty will look "significant" by chance alone. Researchers who scan hundreds of relationships and report only the winners produce exactly the kind of headline that later fails to replicate.
What stronger evidence looks like
Causal claims become convincing when evidence arrives from several directions at once: a plausible mechanism, a relationship that holds across different populations, exposure that comes before the effect, and a dose-response pattern in which more of the cause brings more of the outcome. Epidemiologists used reasoning of this kind to establish that smoking causes lung cancer, long before ethics allowed a direct experiment.
No single correlation can carry that weight on its own. So when a headline says "linked to," read it as "worth investigating."
A better way to read the headline
Ask three questions: how large is the correlation, how many observations support it, and could something else drive both variables? If the story answers none of them, it is a hypothesis, not a finding.