Chapter I
Laplace's Method
The theorem that bears Bayes's name, from probability theory, says how to turn the probability of the data given a cause into the probability of the cause given the data. Pierre-Simon Laplace turned it into a working method. From 1774 he used it to estimate the masses of planets, to decide whether boys are genuinely more likely than girls to be born, and to judge the reliability of witnesses. His rule of succession, derived by assuming that before any evidence every chance is equally likely, became the method's best-known result and its best-known target.
Chapter II
The Eclipse
Critics asked where the prior came from. George Boole and John Venn objected that "equally likely" was an arbitrary assumption dressed up as ignorance, and that a different way of describing the same ignorance gave a different answer. Ronald Fisher agreed and in 1922 built his statistics on the likelihood alone. With the Neyman–Pearson theory of tests from statistical inference, frequentist methods, which speak only of long-run error rates, became the orthodoxy of the twentieth century.
A few people kept the other view alive. Frank Ramsey and Bruno de Finetti showed that anyone whose betting odds break the rules of probability can be made to lose money whatever happens, so rational degrees of belief must be probabilities. Harold Jeffreys, a geophysicist, wrote a Bayesian manual for scientists in 1939, and Leonard Jimmie Savage gave the subject axioms in 1954.
The most consequential Bayesian work was secret. At Bletchley Park, Alan Turing attacked the naval Enigma with a procedure he called Banburismus. Each clue multiplied the odds on a candidate setting, so he worked with logarithms and simply added the scores, in units of decibans. I. J. Good, his assistant, later developed the ideas in public and spent a career arguing for them.
Chapter III
A Closer Look: Will the Sun Rise Tomorrow?
Suppose a coin of unknown bias lands heads 7 times in 10 tosses. What should we believe about its chance of heads?
Laplace started from a flat prior: every value of between 0 and 1 equally plausible. The likelihood of 7 heads and 3 tails is proportional to . Multiplying by the flat prior, the posterior is proportional to , a beta distribution written . In general, starting from and seeing heads and tails gives : updating just adds the counts. The posterior mean is
a little closer to one half than the raw frequency 0.7, because the prior acts like one extra head and one extra tail. The posterior also gives a direct answer to a question the frequentist cannot phrase: the probability that the coin favours heads, , is .
| Evidence | Posterior | Mean | |
|---|---|---|---|
| None | Beta(1, 1) | 0.500 | 0.500 |
| 7 heads, 3 tails | Beta(8, 4) | 0.667 | 0.667 |
| heads, no tails | Beta(+1, 1) |
The last row is the rule of succession. Laplace applied it to the sunrise, assuming, for the sake of the example, that recorded history covered 5,000 years, or 1,826,213 days. The probability of another sunrise comes out as , odds of 1,826,214 to 1, the figure Laplace gave. He added at once that anyone who knows the laws governing the Sun would bet far more heavily. The example was meant to show the method, and critics ever since have used it to mock the flat prior.
Chapter IV
Computing the Posterior
For a coin the posterior has a formula. For a model with hundreds of unknowns it is an integral in hundreds of dimensions that nobody can do. The way out came from physics. In 1953 a Los Alamos team including Nicholas Metropolis and Arianna Rosenbluth, who wrote the program, sampled the states of a simulated liquid with a random walk, the Metropolis algorithm, designed so that it visits each state in proportion to its probability. Averages along the walk then give the answer. In 1990 Alan Gelfand and Adrian Smith showed statisticians that the same trick, in the form of the Gibbs sampler, computes Bayesian posteriors in general. Within a decade Bayesian methods were everywhere, from genetics to cosmology.
The random walks are Markov chains, the subject of stochastic processes, and how long they must run to be trusted is still open. Priors that learn from data, and posteriors over millions of parameters, now connect Bayesian statistics to statistical learning theory.