Chapter I
Combining Observations
Astronomers were the first to face the problem. They had more measurements than unknowns, and the measurements disagreed. Which orbit fits them best? In 1805 Adrien-Marie Legendre proposed choosing the orbit that makes the sum of the squared errors as small as possible. In 1809 Carl Friedrich Gauss published the same method, claimed he had used it for years, and gave it a reason. If errors follow a bell-shaped curve, least squares picks the most probable orbit. Legendre was indignant, and the priority dispute was never settled to both men's satisfaction. Laplace then tied the bell curve to his central limit theorem from probability theory. Errors made of many small causes are normally distributed, so least squares is the right method almost everywhere.
Chapter II
Measuring Variation
For most of the nineteenth century statistics measured the heavens and averaged away variation. Francis Galton, Darwin's cousin, made variation the object of study. In 1886 he found that tall parents have children who are tall, but on average less tall than their parents. He called it regression towards mediocrity and drew a line through the data to measure it. Karl Pearson turned Galton's ideas into mathematics, founded the journal Biometrika, and in 1900 gave the chi-squared test: one number measuring the misfit between observed counts and a theory, with a known distribution when the theory is true.
Pearson's methods assumed large samples. William Sealy Gosset, a brewer at Guinness, had samples of four or five. In 1908, writing as "Student", he found how the average of a small sample really behaves when its spread is estimated from the same data. The t-test is still one of the most used significance tests in science.
Chapter III
Fisher and His Rivals
Ronald Fisher spent fourteen years at the Rothamsted agricultural station, where decades of harvest records had never been analysed properly. From 1922 he rebuilt statistics around the likelihood, the probability of the data as a function of the unknown quantities, and argued that estimating by its maximum is, in large samples, as accurate as any method can be. Later work made this precise. He invented the analysis of variance and insisted that treatments be assigned to plots at random. Randomisation means that the only differences between groups, apart from the treatment, are due to chance, and chance can be calculated.
Jerzy Neyman and Egon Pearson wanted a test to be a rule with guaranteed error rates. In 1933 they framed it as a choice between two hypotheses and found the best tests. Fisher thought this confused science with quality control, and the two sides fought for thirty years. Textbooks later merged their approaches into a single ritual: compute a p-value, compare it with 0.05, declare a result.
Chapter IV
A Closer Look: The Lady Tasting Tea
Fisher's favourite example of an experiment was a colleague at Rothamsted, usually identified as Muriel Bristol, who claimed she could tell whether milk had been poured into the cup before or after the tea. How should the claim be tested?
Fisher's design: prepare eight cups, four of each kind, and present them in random order. She knows there are four of each and must pick the four with milk first. If she cannot taste the difference, every choice of four cups is equally likely. The number of ways to choose four cups from eight is
so the chance of picking all four correctly by luck is , about 1.4%. That is the p-value of a perfect score. Getting three right is much weaker evidence. There are ways to pick exactly three correct cups and one wrong one, so three or more correct happens by chance of the time, about 24%. The design decides in advance how strong the evidence can be.
Pearson's chi-squared test measures the misfit between counts and a theory. Mendel reported 7,324 pea seeds from his hybrids, and his theory predicted round and wrinkled seeds in the ratio 3 to 1:
| Round | Wrinkled | |
|---|---|---|
| Observed | 5,474 | 1,850 |
| Expected (3 : 1) | 5,493 | 1,831 |
Each count is off by 19. The statistic adds up the squared misfits, each scaled by its expected count:
With one degree of freedom, a value this large or larger occurs by chance about 61% of the time, so the data sit comfortably with the theory. A value above 3.84 would have been needed to reject it at the 5% level. In 1936 Fisher applied the same test to all of Mendel's experiments together and found the fit too good: chance would rarely give results so close to the predictions. Whether Mendel, an assistant or simple selective reporting was responsible is still debated.
Chapter V
Significance Under Pressure
By the late twentieth century, decided what journals published and what careers were built on. In 2005 John Ioannidis argued that, given small studies and flexible analyses, most published findings were probably false. Large replication projects in psychology, cancer biology and economics found many famous results that did not repeat. The mathematics of a single test was not wrong. The problem was that a p-value only means what it says if the analysis was fixed before the data were seen and every result was reported.
Statistics answered with more than one proposal. Some want stricter thresholds, some want estimates and intervals instead of verdicts, and some want to measure evidence the Bayesian way, the approach Fisher had tried to banish, which returns in Bayesian statistics. The same questions of fitting and generalising, asked of machines instead of scientists, became statistical learning theory.