Skip to content
Field Atlas

Atlas / Mathematics / The Statistics Thread

Field · Emerged 1774 – 1990

Bayesian Statistics

How should a degree of belief be updated as evidence arrives, and can probability measure belief at all?

4 chapters4 min read6 turning points1 open problem

Branched from
Probability Theory + Statistical Inference
Branched into
Not yet surveyed past here
Figures
Pierre-Simon Laplace, Ronald Fisher, Frank Ramsey, Bruno de Finetti, Leonard Jimmie Savage, Harold Jeffreys, Alan Turing, I. J. Good, Alan Gelfand, Adrian Smith

In brief

Bayesian statistics treats probability as a degree of belief. Anything uncertain, including an unknown constant of nature, gets a probability distribution. Before the data this is the prior. Bayes' theorem combines it with the data to give the posterior, and the posterior is the complete answer: every estimate, interval and prediction is read off from it.

Laplace used the method for fifty years, but in the twentieth century it was nearly banished. Its critics objected that priors are subjective, and the frequentist methods of Fisher and Neyman took over. A small group, Jeffreys, de Finetti, Savage and the code-breakers of Bletchley Park, kept it alive. What finally brought it back was computation: from 1990 Markov chain Monte Carlo made Bayesian answers computable for realistic models, and the approach now runs through science and machine learning.

Key ideas

Prior and posteriorEnters 1774 – 1814

The prior is the distribution of belief before seeing the data. The posterior is the distribution after. Bayes' theorem says the posterior is proportional to the prior times the likelihood of the data.

Subjective probabilityEnters 1926 – 1954

Probability as a coherent degree of belief of a particular person. De Finetti showed that beliefs which violate the rules of probability can be exploited by a series of bets that loses whatever happens.

Weight of evidenceEnters 1940 – 1941

The logarithm of the factor by which evidence multiplies the odds of a hypothesis. Turing measured it in "bans" and tenths of a ban, so that independent clues simply add up.

Markov chain Monte CarloEnters 1990

Computing a posterior by running a random walk whose long-run behaviour is exactly that posterior, and averaging along the walk. It replaced integrals that could not be done by hand with simulations a computer can do.

Objective priorEnters 1939

A prior chosen by a rule instead of personal judgement, meant to express ignorance. Jeffreys's rule gives the same answer however the unknown quantity is measured.

Chapter I

Laplace's Method

The theorem that bears Bayes's name, from probability theory, says how to turn the probability of the data given a cause into the probability of the cause given the data. Pierre-Simon Laplace turned it into a working method. From 1774 he used it to estimate the masses of planets, to decide whether boys are genuinely more likely than girls to be born, and to judge the reliability of witnesses. His rule of succession, derived by assuming that before any evidence every chance is equally likely, became the method's best-known result and its best-known target.

Chapter II

The Eclipse

Critics asked where the prior came from. George Boole and John Venn objected that "equally likely" was an arbitrary assumption dressed up as ignorance, and that a different way of describing the same ignorance gave a different answer. Ronald Fisher agreed and in 1922 built his statistics on the likelihood alone. With the Neyman–Pearson theory of tests from statistical inference, frequentist methods, which speak only of long-run error rates, became the orthodoxy of the twentieth century.

A few people kept the other view alive. Frank Ramsey and Bruno de Finetti showed that anyone whose betting odds break the rules of probability can be made to lose money whatever happens, so rational degrees of belief must be probabilities. Harold Jeffreys, a geophysicist, wrote a Bayesian manual for scientists in 1939, and Leonard Jimmie Savage gave the subject axioms in 1954.

The most consequential Bayesian work was secret. At Bletchley Park, Alan Turing attacked the naval Enigma with a procedure he called Banburismus. Each clue multiplied the odds on a candidate setting, so he worked with logarithms and simply added the scores, in units of decibans. I. J. Good, his assistant, later developed the ideas in public and spent a career arguing for them.

Chapter III

A Closer Look: Will the Sun Rise Tomorrow?

Suppose a coin of unknown bias lands heads 7 times in 10 tosses. What should we believe about its chance pp of heads?

Laplace started from a flat prior: every value of pp between 0 and 1 equally plausible. The likelihood of 7 heads and 3 tails is proportional to p7(1−p)3p^7(1-p)^3. Multiplying by the flat prior, the posterior is proportional to p7(1−p)3p^7(1-p)^3, a beta distribution written Beta(8,4)\text{Beta}(8, 4). In general, starting from Beta(a,b)\text{Beta}(a, b) and seeing hh heads and tt tails gives Beta(a+h,b+t)\text{Beta}(a+h, b+t): updating just adds the counts. The posterior mean is

aa+b=812=23≈0.667,\frac{a}{a+b} = \frac{8}{12} = \frac{2}{3} \approx 0.667 ,

a little closer to one half than the raw frequency 0.7, because the prior acts like one extra head and one extra tail. The posterior also gives a direct answer to a question the frequentist cannot phrase: the probability that the coin favours heads, P(p>1/2)P(p > 1/2), is 1816/2048≈0.891816/2048 \approx 0.89.

EvidencePosteriorMeanP(next is heads)P(\text{next is heads})
NoneBeta(1, 1)0.5000.500
7 heads, 3 tailsBeta(8, 4)0.6670.667
nn heads, no tailsBeta(nn+1, 1)(n+1)/(n+2)(n+1)/(n+2)(n+1)/(n+2)(n+1)/(n+2)

The last row is the rule of succession. Laplace applied it to the sunrise, assuming, for the sake of the example, that recorded history covered 5,000 years, or 1,826,213 days. The probability of another sunrise comes out as 1,826,214/1,826,2151{,}826{,}214 / 1{,}826{,}215, odds of 1,826,214 to 1, the figure Laplace gave. He added at once that anyone who knows the laws governing the Sun would bet far more heavily. The example was meant to show the method, and critics ever since have used it to mock the flat prior.

Chapter IV

Computing the Posterior

For a coin the posterior has a formula. For a model with hundreds of unknowns it is an integral in hundreds of dimensions that nobody can do. The way out came from physics. In 1953 a Los Alamos team including Nicholas Metropolis and Arianna Rosenbluth, who wrote the program, sampled the states of a simulated liquid with a random walk, the Metropolis algorithm, designed so that it visits each state in proportion to its probability. Averages along the walk then give the answer. In 1990 Alan Gelfand and Adrian Smith showed statisticians that the same trick, in the form of the Gibbs sampler, computes Bayesian posteriors in general. Within a decade Bayesian methods were everywhere, from genetics to cosmology.

The random walks are Markov chains, the subject of stochastic processes, and how long they must run to be trusted is still open. Priors that learn from data, and posteriors over millions of parameters, now connect Bayesian statistics to statistical learning theory.

Applications

Where it is used

  • Search and rescue

    Finding what is lost

    Bayesian search theory divides the sea into cells, gives each a prior probability, and updates after every unsuccessful search. It guided the hunt for a lost hydrogen bomb off Palomares in 1966 and, in 2011, the search for Air France flight 447, found within a week of resuming the search in the area the analysis ranked highest.

    › Sources (1)
    • Stone, L. D., Keller, C. M., Kratzke, T. M. & Strumpfer, J. P. (2014). Search for the wreckage of Air France Flight AF 447. Statistical Science 29(1): 69–80.
  • Evolution↗ Biology · Phylogenetics

    Bayesian family trees

    Reconstructing the tree of life from DNA means choosing among astronomically many possible trees. Bayesian phylogenetics samples trees by MCMC in proportion to their posterior probability, and reports how certain each branch is.

    › Sources (1)
    • Huelsenbeck, J. P. & Ronquist, F. (2001). MRBAYES: Bayesian inference of phylogenetic trees. Bioinformatics 17(8): 754–755.
  • Cosmology↗ Physics · Physical Cosmology

    Weighing the universe

    The amounts of ordinary matter, dark matter and dark energy are estimated by fitting models to the cosmic microwave background with Markov chain Monte Carlo, which maps out the posterior of half a dozen parameters at once.

    › Sources (1)
    • Lewis, A. & Bridle, S. (2002). Cosmological parameters from CMB and other data: a Monte Carlo approach. Physical Review D 66: 103511.

Open problems

Where the map runs out

Open

How long must a sampler run?

Open in general as of 2026; sharp answers exist only for special classes of chains.

A Markov chain Monte Carlo sampler is only correct in the long run. How many steps are needed before its output is a fair sample from the posterior? In practice, convergence is judged by diagnostics that can detect some failures but never prove success.

Why it is hard

The mixing time depends on the shape of the posterior in spaces with thousands or millions of dimensions, which is exactly what is unknown. Rigorous bounds, from the geometry and spectra of Markov chains, exist for card shuffling and some physical models, but rarely for the models statisticians actually fit.

What resolving it unlocks

Guaranteed error bars on Bayesian answers, and principled ways to design faster samplers for the large models of genetics, cosmology and machine learning.

› Sources (1)
  • Diaconis, P. (2009). The Markov chain Monte Carlo revolution. Bulletin of the American Mathematical Society 46(2): 179–205.

Further reading

  1. McGrayne, S. B. (2011). The Theory That Would Not Die. Yale University Press.

    A popular history of Bayes' rule, from Laplace to the code-breakers and beyond.

  2. Jaynes, E. T. (2003). Probability Theory: The Logic of Science. Cambridge University Press.

    A forceful argument that probability is the logic of plausible reasoning.

  3. Gelman, A., Carlin, J. B., Stern, H. S., Dunson, D. B., Vehtari, A. & Rubin, D. B. (2013). Bayesian Data Analysis (3rd ed.). CRC Press.

    The standard modern textbook.