Experiments & Hypothesis Testing
The Blind Taste Test: Can You Really Tell the Difference?
Two cups, identical to the eye. One holds your favorite brand, the other a cheaper rival. Someone shuffles them and you take a sip. Blind tests are a delight at parties and a genuine tool in food science, but interpreting the results takes more than counting hands.

What chance alone looks like
Suppose each taster must pick the cup they prefer, and there is truly no difference. Each answer is then a coin flip with p = 0.5. If 7 of 10 tasters choose brand A, that feels like a verdict, yet the binomial distribution calculator shows a 17% chance of seeing 7 or more by luck alone. With 15 of 20, the chance drops to about 2.1%, and the result starts to carry weight.
The one-sample proportion test formalizes this: set the null proportion at 0.5, enter the counts, and read the p-value. Exact binomial calculations are preferable with small groups, since the normal approximation becomes unreliable.
Remember that failing to reach significance does not prove tasters cannot tell the difference. It means the test was not strong enough to show it. With a small group, "no evidence of a difference" and "evidence of no difference" are very different statements, so report the observed counts alongside the verdict.
The triangle test
Food scientists often prefer a triangle test: three cups, two identical and one different, and each taster must spot the odd one out. Pure guessing then succeeds one time in three. If 11 of 20 tasters pick correctly, the chance of that many or more under guessing is about 3.8%. With 12 of 20 it falls to 1.3%, and 11 is the first count at which a group of 20 clears the usual 5% bar.
Notice how the baseline matters. Eleven correct answers out of twenty looks unremarkable in a two-cup test, where chance already predicts 10, but it is meaningful in a triangle test, where chance predicts fewer than seven. Professional sensory panels train their tasters and use many repetitions for exactly this reason.
How big is the effect?
A p-value says whether guessing is plausible, not how good the tasters are. With 15 of 20 correct, the sample proportion is 75%, and a simple confidence interval for a single proportion runs from about 56% to 94%. That is wide, a reminder that twenty sips cannot pin down skill precisely.
Now imagine three brands and 40 tasters choosing a favorite: 18, 12 and 10 votes. If no brand were better, each would expect about 13.3 votes. The chi-square goodness-of-fit calculator gives a p-value near 0.27, so the apparent winner is well within normal noise.
Design choices that decide everything
Sample size sets your power. With only ten tasters, even a real preference may fail to reach significance. Detecting a true 70% preference reliably takes roughly 40 to 50 tasters, far more than most informal tastings recruit.
Independence matters as well. Ten sips from one person amount to one taster's opinion, not ten. Decide the number of tasters beforehand, because checking the score after each person and stopping once it looks significant inflates false positives, the same trap that catches product teams in A/B testing.
Finally, comparing several brands pairwise multiplies the chances of a lucky difference. The more comparisons you run, the more cautious each one deserves to be.
Good tastings also randomize the order, hide labels and use identical cups, because expectations and position influence what people report.
A statistic can only be as honest as the experiment behind it.