Chapter I
Counting the Dead
London's parishes published a weekly bill of mortality: how many were buried, and of what. It existed as an early-warning system for plague, and nobody had tried to reason from it until John Graunt, a haberdasher with a taste for arithmetic, worked through several decades of them and published his conclusions in 1662.
He estimated the population of the city, which no one knew. He noticed that slightly more boys than girls are born, consistently, year after year. He separated causes of death that occur at a steady rate from those that come in waves, which is the distinction between endemic and epidemic. And he constructed a table showing how many of a hundred people born would be expected to survive to each decade — the first life table, the ancestor of every actuarial calculation since.
Nearly two centuries later William Farr made the counting official. Appointed to the new General Register Office, he imposed a classification of causes of death, and insisted on two things that look obvious and are not: that counts be divided by the population at risk, and that comparisons between places be standardised for age, since a district full of young people will have fewer deaths whatever its conditions. His analysis of the 1849 cholera epidemic found that mortality fell steadily with height above the Thames, which he attributed to bad air being denser at low elevation. The pattern was real. The explanation was wrong, and the right one — that elevation predicts whose water came from the river — was being assembled a few streets away by John Snow, whose investigation is a waypoint on microbiology.
That is the characteristic difficulty of the whole field in one example. An association can be solid and repeatable and still point to the wrong cause.
Chapter II
Making the Groups Comparable
The only clean way to compare two groups of people is to decide which is which by chance. Austin Bradford Hill built the first properly randomised therapeutic trial around a shortage: the Medical Research Council had a little streptomycin and many patients with tuberculosis. Allocation was by sealed envelopes in a predetermined random order, controls received bed rest alone, and the chest X-rays were read by assessors who did not know which patients had received the drug. Four of 55 treated patients died within six months, against 14 of 52 controls.
Randomisation is unavailable for most of what makes people ill. Nobody can be assigned to smoke for thirty years. The two designs that work without it both have a characteristic weakness. A case-control study collects people who already have the disease and compares their reported exposures with those of controls, which is quick and cheap and depends entirely on whether the controls were drawn from the same population as the cases. A cohort study enrols people before anyone is ill and follows them, which gives absolute risks and takes decades.
Richard Doll and Bradford Hill used both on the same question. In 1950, lung cancer deaths in Britain had risen fifteenfold in three decades, and the leading suspects were tarred roads and car exhaust. Their case-control study of 649 male patients found that almost none of them were non-smokers. Doubting their own design, they then recruited 40,000 doctors and followed them forward for decades, which produced the two findings that settled the matter: risk rose with the number of cigarettes, and it fell after quitting.
Ronald Fisher, who had done more than anyone to establish randomisation as a principle, spent a decade arguing that they were wrong — that some constitutional factor might cause both the taste for tobacco and the susceptibility to cancer. He was mistaken, and the objection was exactly the right shape. Confounding by something unmeasured is the permanent occupational hazard of this field, and Fisher's error was not in raising it but in refusing the evidence that answered it.
Chapter III
A Closer Look: What 649 Interviews Established
Doll and Hill's 1950 study is small enough to work through. Of 649 male lung-cancer patients, 0.3% were non-smokers — two men. Of 649 matched hospital controls, 4.2% were non-smokers, or 27 men. The four cells are:
| Smokers | Non-smokers | |
|---|---|---|
| Lung cancer cases | 647 | 2 |
| Controls | 622 | 27 |
A case-control study cannot give the risk of cancer among smokers, because the number of cases was fixed by the investigators. What it gives is the odds ratio:
The odds of having smoked are fourteen times higher among the cases. For a disease this rare in the population, the odds ratio approximates the relative risk, so smokers are of the order of fourteen times more likely to develop lung cancer.
Now notice what that figure does not say. The lifetime risk of lung cancer for a lifelong non-smoker is roughly 0.5%. Multiplying by 14 gives about 7% — a large and terrible number, and also a reminder that most smokers do not get lung cancer, which was a common argument against the finding at the time. A fourteenfold relative risk on a small base is still a small absolute risk for any one person, and an enormous one for a country where most men smoked.
The strength of the association is what makes it hard to explain away, and this is the first of Bradford Hill's nine considerations. Suppose Fisher were right that some genetic factor causes both smoking and cancer. For that confounder to manufacture a fourteenfold association, it would have to be strongly associated with both — roughly, if it multiplied cancer risk by and the odds of smoking by , the spurious association is bounded by something like the smaller of and . Producing 14 from nothing requires a hidden factor with an effect on cancer larger than almost any known risk factor, and an equally strong effect on behaviour. Possible; implausible; and testable, because it predicts no dose-response and no benefit from quitting.
The cohort study delivered both tests. Doctors smoking 25 or more cigarettes a day had about 20 to 25 times the lung-cancer mortality of non-smokers; those smoking a few had two or three times. Doctors who stopped had risks that declined with years since quitting, approaching but not reaching that of those who never smoked. A constitutional confounder does not behave that way. Changing behaviour changed outcome, which is as close to an experiment as the data could come.
This is the pattern of a successful epidemiological argument: a large effect, a gradient, a reversal on removal of exposure, and a mechanism that fits. Where only the first is available — a modest association, no gradient, no intervention — the record of the field is much worse, which is what the open problem below is about.
Chapter IV
Nine Considerations and Their Misuse
In 1965 Bradford Hill set out what he had learned about the inference. Strength of association, consistency across studies and populations, specificity of effect, temporality — exposure must precede disease — a biological gradient, plausibility of mechanism, coherence with other knowledge, experimental evidence where available, and analogy with similar established causes.
He said plainly that these are not a checklist and cannot be scored, that none except temporality is required, and that the real question is always whether some alternative explanation is more likely than cause and effect. They have been used as a scorecard ever since.
The field's subsequent history has justified his caution. Hormone replacement therapy was observed to protect against heart disease across many cohorts, and the randomised trial found the opposite. Beta-carotene was observed to protect against lung cancer, and the trial found increased incidence among smokers. Both reversals happened because the exposed groups differed systematically from the unexposed in ways the adjustments did not capture, and both are the reason that where randomisation is possible, nothing else is accepted.
Where it is impossible, the field has turned to designs that approximate it: natural experiments, instrumental variables, and Mendelian randomisation, which uses the random allocation of genetic variants at conception as a proxy for assigning an exposure. The methods are increasingly statistical, and the questions increasingly arrive from the laboratory — from the mechanism of cancer, the transmission of infection, and the effects of the drugs that pharmacology produces.