Statistical Calculators

Medicine & Health Statistics

How to Read a Clinical Trial: Odds Ratios, P-Values and the Headline Trap

Published · 3 min read

A new treatment is announced with a bold figure: "80% effective." Another study reports a "statistically significant improvement" that patients barely notice. Both headlines can be accurate and still mislead. Knowing what sits behind such claims is one of the most practical forms of science literacy.

Illustration of a clinical trial comparing treatment and placebo groups, with odds ratio and p-value statistics

Start with the two-by-two table

Nearly every trial boils down to four counts: treated and untreated people, with and without the outcome. Imagine a hypothetical trial where 8 of 10,000 vaccinated volunteers fall ill, against 40 of 10,000 on placebo. Risk drops from 0.40% to 0.08%, which is the famous 80% efficacy. The odds ratio calculator returns about 0.20, the same story in a form researchers can combine across studies.

One caution: relative numbers can flatter. A drop from 0.40% to 0.08% is an 80% relative reduction, but the absolute difference is only 0.32 percentage points. Both are true, and both belong in an honest summary.

Is it chance? Choose the right test

To check whether such a gap could be luck, a chi-square test of independence compares observed counts with those expected if the treatment did nothing. It behaves well when expected counts are not tiny.

In small early-phase trials, say 3 recoveries among 20 patients versus 9 among 20, the Fisher exact test is safer, because it computes the probability directly instead of relying on an approximation. And when the same patients are assessed before and after treatment, McNemar's test respects the pairing.

Significant is not the same as important

A p-value below 0.05 only says the result would be unusual if the treatment did nothing. It says nothing about size, and it is not the probability that the treatment works.

A giant trial can flag a trivial benefit as significant, while a small one can miss a real effect entirely.

So always look for three things beside the p-value: the effect size, a confidence interval, and the number of people the estimate rests on. A wide interval from a few dozen patients is a hint, not a verdict.

How certain is that 0.20?

A single odds ratio is a best guess. Using the hypothetical counts above, the 95% confidence interval runs from roughly 0.09 to 0.43. Even the cautious end of that range shows a large benefit, which is far more reassuring than a point estimate alone.

Another useful translation is the number needed to treat. The absolute risk fell by 0.32 percentage points, so about 313 people would need the vaccine during the trial period to prevent one case. Such figures help patients and policymakers weigh benefits against side effects and costs.

Finally, check the design. Was allocation randomized, were participants and doctors blinded, and was the main outcome declared before the data came in? Strong statistics cannot rescue a weak design, but a strong design makes the statistics worth trusting.

Beware of subgroup claims too. When results are sliced by age, sex and region, some slices will look striking purely by chance. A finding from the pre-specified main analysis deserves more trust than one discovered afterwards. Likewise, check that the outcome reflects something patients care about, such as hospitalization or survival, rather than a lab marker that only hints at it.

A five-minute habit

Next time a headline promises a breakthrough, try to rebuild its two-by-two table. If the article gives you too little to do that, treat the claim with caution. If it does, the calculators will tell you in seconds whether the excitement is justified.

Check your own counts: start with the odds ratio calculator or the Fisher exact test and see how conclusions change when the sample is small.