Skip to content
Field Atlas

Atlas / Mathematics / The Analysis Thread

Field · Emerged 1654 – 1933

Probability Theory

How can chance be measured, and what laws does randomness obey in the long run?

5 chapters3 min read5 turning points1 open problem

Branched from
Calculus + Real Analysis
Branched into
Bayesian Statistics + Game Theory + Information Theory + Monte Carlo Methods + Probabilistic Combinatorics + Statistical Inference + Stochastic Processes
Figures
Pierre de Fermat, Blaise Pascal, Jacob Bernoulli, Abraham de Moivre, Pierre-Simon Laplace, Thomas Bayes, Richard Price, Andrey Kolmogorov

In brief

Probability theory is the mathematics of uncertainty. It assigns numbers between 0 and 1 to events and works out what follows. Individual random events are unpredictable, but in bulk they obey firm laws: averages settle down (the law of large numbers), and sums of many small random effects take a bell shape (the central limit theorem).

It began with gamblers' questions in the seventeenth century and was long regarded as not quite respectable mathematics. In 1933 Kolmogorov founded it on measure theory, from real analysis, and it became the basis of statistics, statistical physics, genetics, finance and machine learning.

Key ideas

Probability spaceEnters 1933

The set of possible outcomes, a collection of events, and a measure assigning each event a probability, with total probability 1. Kolmogorov's framework: probability is measure theory.

Expected valueEnters 1654

The long-run average of a random quantity, each outcome weighted by its probability. The idea grew from Pascal and Fermat's solution to the problem of points, and Huygens made it explicit in 1657.

Law of large numbersEnters 1713

The average of many independent repetitions converges to the expected value. It is why casinos and insurers can count on their long-run averages.

Central limit theoremEnters 1733 – 1810

Sums of many independent random effects are approximately normally distributed, whatever each effect looks like. It explains why the bell curve is everywhere.

Bayes' theoremEnters 1763 – 1774

P(A∣B)=P(B∣A) P(A)/P(B)P(A \mid B) = P(B \mid A)\,P(A) / P(B): how to update the probability of a hypothesis in light of evidence. It is the basis of Bayesian statistics and much of machine learning.

Draws on other domains

Chapter I

A Gambler's Question

In 1654 the Chevalier de Méré, a gambler, put a puzzle to Blaise Pascal. If a game of chance is interrupted before anyone has won, how should the stakes be divided? Pascal wrote to Pierre de Fermat, and in a few letters they solved it. Count every way the game could have continued and divide in proportion. For the first time, uncertainty was a quantity that could be calculated. Christiaan Huygens wrote the first textbook on the subject three years later.

Chapter II

Laws of Large Numbers

Chance individually is unpredictable, yet it has laws in bulk. Jacob Bernoulli spent twenty years proving the first. As an experiment is repeated, the fraction of successes converges to the true probability. His Ars Conjectandi appeared posthumously in 1713. Abraham de Moivre, a French Protestant exile in London who earned his living partly advising gamblers, found in 1733 the shape of the fluctuations, the bell curve. Pierre-Simon Laplace generalised it to the central limit theorem: sums of many independent small effects are always approximately normal, whatever the effects. It explains why heights, measurement errors and exam scores so often follow the same curve.

Reasoning backwards, from observed data to the chance behind it, came from Thomas Bayes, whose essay was published by a friend after his death in 1763, and more powerfully from Laplace. Their "inverse probability" is today's Bayesian inference.

Chapter III

Respectability

Through the nineteenth century probability was useful but suspect. Its founding notions ("equally likely", "at random") were circular, and paradoxes showed that different, equally natural ways of choosing "at random" gave different answers. Hilbert listed its foundations, within his sixth problem, among the major problems of 1900.

The solution came from real analysis. In 1933 Andrey Kolmogorov observed that Lebesgue's measure theory already had exactly the right structure. Probability is a measure of total size 1 on a space of outcomes, events are measurable sets, and expectation is the Lebesgue integral. The paradoxes dissolved into precise statements, and probability became a full branch of mathematics.

Chapter IV

A Closer Look: The Test That Is 99% Accurate

A disease affects 1 in 100 people. A test detects it 99% of the time when it is present, and gives a false positive 5% of the time when it is not. You test positive. What is the chance you have the disease?

Most people guess about 95%. Bayes' theorem gives the answer. Picture 10,000 people. About 100 have the disease, and 99 of them test positive. Of the 9,900 healthy people, 5% also test positive: 495 of them. So there are 99+495=59499 + 495 = 594 positive results, of which only 99 are true:

P(disease∣positive)=0.99×0.010.99×0.01+0.05×0.99=99594=16≈17%.P(\text{disease} \mid \text{positive}) = \frac{0.99 \times 0.01}{0.99 \times 0.01 + 0.05 \times 0.99} = \frac{99}{594} = \frac{1}{6} \approx 17\% .

A positive result raises the probability from 1% to about 17%, a big jump, but most positives are still false alarms, because healthy people vastly outnumber sick ones. This is why screening programmes follow a positive result with a second, independent test. If the second test is also positive, Bayes' theorem applied again, starting from 17%, gives about 80%.

The same reasoning, updating a probability as evidence arrives, runs spam filters, medical diagnosis, forensic statistics and much of machine learning. Its misuse has consequences as well. Confusing P(evidence∣innocent)P(\text{evidence} \mid \text{innocent}) with P(innocent∣evidence)P(\text{innocent} \mid \text{evidence}), the "prosecutor's fallacy", has contributed to wrongful convictions.

Chapter V

Everywhere

Probability now runs through science. It is the mathematics of genetic drift in population genetics, of Brownian motion and statistical physics, of statistics, finance and machine learning. Its frontier includes random structures whose behaviour at a critical point, where a sudden global change happens, is still out of reach, most famously percolation in three dimensions.

Applications

Where it is used

  • Population genetics↗ Biology · Population Genetics

    Genetic drift is a random walk

    In a finite population, which individuals happen to reproduce is partly chance, so allele frequencies wander randomly. The Wright–Fisher model and its diffusion approximations, pure probability theory, are how population genetics measures drift and dates evolutionary events.

    › Sources (1)
    • Ewens, W. J. (2004). Mathematical Population Genetics (2nd ed.). Springer.
  • Statistical physics↗ Physics · Kinetic Theory of Gases

    Brownian motion and atoms

    In 1905 Einstein explained the jittering of tiny particles suspended in water as the result of random molecular collisions, and predicted how far they would wander. Perrin's measurements confirmed it, convincing sceptics that atoms exist. Random walks are now a basic tool of physics.

    › Sources (1)
    • Einstein, A. (1905). Über die von der molekularkinetischen Theorie der Wärme geforderte Bewegung von in ruhenden Flüssigkeiten suspendierten Teilchen. Annalen der Physik 17: 549–560.
  • Finance

    Pricing risk

    Insurance premiums rest on the law of large numbers, and modern derivative pricing on random walks. The Black–Scholes model of 1973 priced options by treating stock prices as following a form of Brownian motion.

    › Sources (1)
    • Black, F. & Scholes, M. (1973). The pricing of options and corporate liabilities. Journal of Political Economy 81(3): 637–654.
  • Soft matter↗ Physics · Soft Matter

    A polymer is a random walk

    Treating a long molecule as a sequence of randomly oriented steps gives its size as bNb\sqrt{N} rather than bNbN, and the Gaussian distribution of end-to-end distances supplies the entropy from which rubber's elasticity follows — a restoring force proportional to temperature, with no bond being stretched. The excluded-volume correction, which forbids the walk from crossing itself, changes the exponent and was computed by borrowing the renormalisation group from critical phenomena.

    › Sources (2)
    • de Gennes, P.-G. (1979). Scaling Concepts in Polymer Physics. Cornell University Press.
    • Rubinstein, M. & Colby, R. H. (2003). Polymer Physics. Oxford University Press.

Open problems

Where the map runs out

Open

Is 3D percolation continuous at its critical point?

Proved in two dimensions and in high dimensions; open in three dimensions as of 2026.

Randomly keep each edge of a three-dimensional grid with probability pp. Above a critical value, an infinite connected cluster appears. Does an infinite cluster already exist at the critical point? It is believed not, so that the transition is continuous.

Why it is hard

In two dimensions, symmetry and conformal invariance give powerful tools. In high dimensions, a technique called the lace expansion works. Three dimensions falls in between, and none of the known methods reaches it.

What resolving it unlocks

Percolation models how fluids seep through rock, how epidemics spread and how networks fail. Settling the three-dimensional critical behaviour would confirm the physicists' picture of phase transitions in the dimension we live in.

› Sources (1)
  • Duminil-Copin, H. (2018). Sixty years of percolation. Proceedings of the International Congress of Mathematicians 2018: 2829–2856.

Further reading

  1. Devlin, K. (2008). The Unfinished Game: Pascal, Fermat, and the Seventeenth-Century Letter that Made the World Modern. Basic Books.

    The story of the letters that founded probability, for general readers.

  2. Hacking, I. (1975). The Emergence of Probability. Cambridge University Press.

    A philosopher's history of why probability emerged when it did.

  3. Feller, W. (1968). An Introduction to Probability Theory and Its Applications, Vol. 1 (3rd ed.). Wiley.

    The classic textbook, rich in examples, and still widely used.