Skip to content
Field Atlas

Atlas / Mathematics / The Statistics Thread

Field · Emerged 1805 – 1935

Statistical Inference

How can conclusions about a whole population be drawn from a limited sample, with a known risk of being wrong?

5 chapters5 min read7 turning points1 open problem

Branched from
Probability Theory
Branched into
Bayesian Statistics + Statistical Learning Theory
Figures
Carl Friedrich Gauss, Adrien-Marie Legendre, Francis Galton, Karl Pearson, William Sealy Gosset, Ronald Fisher, Jerzy Neyman, Egon Pearson, John Ioannidis

In brief

Statistical inference is probability run backwards. Probability theory starts from a known chance mechanism and predicts the data. Inference starts from the data and asks what mechanism produced them: what the true average is, whether a treatment works, whether two quantities are related. Its answers always carry a margin of error, and its central achievement is to make that margin exact.

It grew from astronomy, where observations had to be combined, through the study of heredity, where Galton and Pearson measured variation, to agriculture, where Fisher invented the randomised experiment. By 1935 it had the tools still taught today: estimates, confidence intervals, significance tests and p-values. Its success made it the gatekeeper of published science, and the replication crisis of the 2010s showed how badly those tools can be misused.

Key ideas

Least squaresEnters 1805 – 1809

Choose the values that make the sum of the squared errors as small as possible. It is the standard way to fit a line or a curve to noisy measurements.

RegressionEnters 1886

Predicting one quantity from another. Galton noticed that extreme parents have children closer to the average, "regression to the mean", and the name stuck to the whole method.

Significance test and p-valueEnters 1900

The p-value is the probability, if nothing but chance were at work, of seeing data at least as extreme as those observed. A small p-value is evidence against pure chance, not the probability that chance is the explanation.

Maximum likelihoodEnters 1922 – 1935

Estimate an unknown quantity by the value that makes the observed data most probable. Fisher argued that in large samples no other method does better, a claim later work made precise.

RandomisationEnters 1922 – 1935

Assign treatments by chance, so that no hidden factor can systematically favour one group. It turns an experiment into a chance mechanism whose behaviour is known exactly.

Draws on other domains

Chapter I

Combining Observations

Astronomers were the first to face the problem. They had more measurements than unknowns, and the measurements disagreed. Which orbit fits them best? In 1805 Adrien-Marie Legendre proposed choosing the orbit that makes the sum of the squared errors as small as possible. In 1809 Carl Friedrich Gauss published the same method, claimed he had used it for years, and gave it a reason. If errors follow a bell-shaped curve, least squares picks the most probable orbit. Legendre was indignant, and the priority dispute was never settled to both men's satisfaction. Laplace then tied the bell curve to his central limit theorem from probability theory. Errors made of many small causes are normally distributed, so least squares is the right method almost everywhere.

Chapter II

Measuring Variation

For most of the nineteenth century statistics measured the heavens and averaged away variation. Francis Galton, Darwin's cousin, made variation the object of study. In 1886 he found that tall parents have children who are tall, but on average less tall than their parents. He called it regression towards mediocrity and drew a line through the data to measure it. Karl Pearson turned Galton's ideas into mathematics, founded the journal Biometrika, and in 1900 gave the chi-squared test: one number measuring the misfit between observed counts and a theory, with a known distribution when the theory is true.

Pearson's methods assumed large samples. William Sealy Gosset, a brewer at Guinness, had samples of four or five. In 1908, writing as "Student", he found how the average of a small sample really behaves when its spread is estimated from the same data. The t-test is still one of the most used significance tests in science.

Chapter III

Fisher and His Rivals

Ronald Fisher spent fourteen years at the Rothamsted agricultural station, where decades of harvest records had never been analysed properly. From 1922 he rebuilt statistics around the likelihood, the probability of the data as a function of the unknown quantities, and argued that estimating by its maximum is, in large samples, as accurate as any method can be. Later work made this precise. He invented the analysis of variance and insisted that treatments be assigned to plots at random. Randomisation means that the only differences between groups, apart from the treatment, are due to chance, and chance can be calculated.

Jerzy Neyman and Egon Pearson wanted a test to be a rule with guaranteed error rates. In 1933 they framed it as a choice between two hypotheses and found the best tests. Fisher thought this confused science with quality control, and the two sides fought for thirty years. Textbooks later merged their approaches into a single ritual: compute a p-value, compare it with 0.05, declare a result.

Chapter IV

A Closer Look: The Lady Tasting Tea

Fisher's favourite example of an experiment was a colleague at Rothamsted, usually identified as Muriel Bristol, who claimed she could tell whether milk had been poured into the cup before or after the tea. How should the claim be tested?

Fisher's design: prepare eight cups, four of each kind, and present them in random order. She knows there are four of each and must pick the four with milk first. If she cannot taste the difference, every choice of four cups is equally likely. The number of ways to choose four cups from eight is

(84)=70,\binom{8}{4} = 70 ,

so the chance of picking all four correctly by luck is 1/701/70, about 1.4%. That is the p-value of a perfect score. Getting three right is much weaker evidence. There are 4×4=164 \times 4 = 16 ways to pick exactly three correct cups and one wrong one, so three or more correct happens by chance 17/7017/70 of the time, about 24%. The design decides in advance how strong the evidence can be.

Pearson's chi-squared test measures the misfit between counts and a theory. Mendel reported 7,324 pea seeds from his hybrids, and his theory predicted round and wrinkled seeds in the ratio 3 to 1:

RoundWrinkled
Observed5,4741,850
Expected (3 : 1)5,4931,831

Each count is off by 19. The statistic adds up the squared misfits, each scaled by its expected count:

χ2=1925493+1921831≈0.066+0.197=0.263.\chi^2 = \frac{19^2}{5493} + \frac{19^2}{1831} \approx 0.066 + 0.197 = 0.263 .

With one degree of freedom, a value this large or larger occurs by chance about 61% of the time, so the data sit comfortably with the theory. A value above 3.84 would have been needed to reject it at the 5% level. In 1936 Fisher applied the same test to all of Mendel's experiments together and found the fit too good: chance would rarely give results so close to the predictions. Whether Mendel, an assistant or simple selective reporting was responsible is still debated.

Chapter V

Significance Under Pressure

By the late twentieth century, p<0.05p < 0.05 decided what journals published and what careers were built on. In 2005 John Ioannidis argued that, given small studies and flexible analyses, most published findings were probably false. Large replication projects in psychology, cancer biology and economics found many famous results that did not repeat. The mathematics of a single test was not wrong. The problem was that a p-value only means what it says if the analysis was fixed before the data were seen and every result was reported.

Statistics answered with more than one proposal. Some want stricter thresholds, some want estimates and intervals instead of verdicts, and some want to measure evidence the Bayesian way, the approach Fisher had tried to banish, which returns in Bayesian statistics. The same questions of fitting and generalising, asked of machines instead of scientists, became statistical learning theory.

Applications

Where it is used

  • Medicine↗ Biology · Epidemiology

    The randomised controlled trial

    Fisher's randomised field plots became the randomised clinical trial. The Medical Research Council's 1948 trial of streptomycin for tuberculosis, designed by Austin Bradford Hill, allocated patients by chance. It is now the standard of evidence for every new drug.

    › Sources (1)
    • Medical Research Council (1948). Streptomycin treatment of pulmonary tuberculosis. British Medical Journal 2(4582): 769–782.
  • Genetics↗ Biology · Population Genetics

    Splitting variation into its causes

    Fisher invented the analysis of variance to divide the variation in a trait among its causes, first to reconcile Mendel with continuous traits. Heritability, the share of variation due to genes, is still estimated this way.

    › Sources (1)
    • Fisher, R. A. (1918). The correlation between relatives on the supposition of Mendelian inheritance. Transactions of the Royal Society of Edinburgh 52: 399–433.
  • Genomics↗ Biology · Genomics

    A million tests at once

    A genome-wide association study tests around a million genetic variants for a link with a disease. At the usual 5% level, tens of thousands would pass by chance, so the field adopted a threshold of p<5×10−8p < 5 \times 10^{-8}, a correction for multiple testing that made its findings replicate.

    › Sources (1)
    • Risch, N. & Merikangas, K. (1996). The future of genetic studies of complex human diseases. Science 273(5281): 1516–1517.

Open problems

Where the map runs out

Open

Making published findings reliable

Open as of 2026; proposed remedies are being tried, and none is agreed as sufficient.

How should evidence from data be summarised and reported so that published findings replicate at the rate their stated error rates promise? Proposals include a stricter threshold of p<0.005p < 0.005, reporting estimates and intervals instead of verdicts, Bayesian measures of evidence, and registering the analysis before the data are seen.

Why it is hard

The mathematics of a single test is settled. The difficulty is that a p-value is only valid if the analysis was fixed in advance and every result is reported, and real research involves many choices made after seeing the data. Measuring the effect of those hidden choices, and designing rules that researchers will follow, is as much a question about incentives as about probability.

What resolving it unlocks

Trustworthy medicine, psychology, economics and biology. Every field that tests hypotheses on noisy data depends on its published record being a fair sample of what was found.

› Sources (2)
  • Wasserstein, R. L. & Lazar, N. A. (2016). The ASA statement on p-values: context, process, and purpose. American Statistician 70(2): 129–133.
  • Benjamin, D. J. et al. (2018). Redefine statistical significance. Nature Human Behaviour 2: 6–10.

Further reading

  1. Stigler, S. M. (1986). The History of Statistics: The Measurement of Uncertainty before 1900. Harvard University Press.

    The standard history, from least squares to Pearson.

  2. Salsburg, D. (2001). The Lady Tasting Tea: How Statistics Revolutionized Science in the Twentieth Century. W. H. Freeman.

    A readable account of the people who built modern statistics.

  3. Stigler, S. M. (2016). The Seven Pillars of Statistical Wisdom. Harvard University Press.

    A short book on the handful of ideas that make statistics work.