Chapter I
Only Novelties Group
By the 1950s everyone agreed that classification should reflect descent. Nobody could say how to get from characters to descent without judgement. Willi Hennig, a specialist in flies who drafted his theory in an Allied prisoner-of-war camp, supplied the missing rule in 1950, and it is almost embarrassingly simple: a shared character groups organisms only if it is new.
Birds and crocodiles both lay shelled eggs, but so did their remote ancestors, and so do turtles and lizards; the character cannot separate any group from any other. Birds and crocodiles also share a four-chambered heart and a particular arrangement of skull openings that their common ancestor acquired and that lizards never had. Those are evidence. Hennig called a shared novelty a synapomorphy and a retained ancestral state a plesiomorphy, and insisted that only the former counts. From the rule follows a severe consequence: a named group must contain all of its ancestor's descendants. Reptilia without birds fails that test and is therefore not a group at all — which is why, in every modern classification, birds are reptiles.
Published in German, the book went unread for sixteen years. Its 1966 translation arrived at the moment when systematics was tearing itself apart over method, and it won, comprehensively and acrimoniously.
Chapter II
Molecules as Documents
The second idea came from chemistry. Emile Zuckerkandl and Linus Pauling lined up haemoglobin sequences from several mammals and counted differences. Horse and human differ in about 18 of 141 residues in the alpha chain; the gorilla differs from the human in one. Plotted against divergence times estimated from fossils, the counts fell roughly on a line. In 1965 they drew the conclusion: a protein is a document of evolutionary history, and its rate of change is a clock that runs whether or not anything visible is happening to the organism.
Morphologists were unimpressed, and the claim was overstated — rates differ between genes, between lineages and between sites, and modelling that variation is still the hardest part of molecular dating. But the consequence was immediate and permanent. Every organism carries a record of its ancestry in a form that can be read without fossils, and the record is the same kind of data for a bacterium, an oak and a whale. Relationships that anatomy could never settle, because the organisms have no anatomy in common, became answerable.
Chapter III
From Counting Steps to Fitting Models
The first computational methods counted. Joseph Camin and Robert Sokal proposed in 1965 that the best tree is the one needing fewest character changes, and tested the idea on plasticine "caminalcules" whose true genealogy they had invented and therefore knew. Margaret Dayhoff and Richard Eck published the first protein tree computed by that principle in 1966. Walter Fitch and Emanuel Margoliash built trees from distance matrices in 1967, and Fitch's 1971 algorithm scored a tree against an alignment in a single pass over the sites.
Then Joseph Felsenstein showed that counting can fail in a specific and dangerous way. If two lineages on opposite sides of the tree both evolve fast, some of their changes will coincide by chance, and parsimony reads coincidence as shared novelty. In 1978 he proved that for some branch lengths the method is statistically inconsistent: more data make it more certain of the wrong tree. The fix was to model substitution explicitly, which he did in 1981 — a tree's likelihood is the probability of the observed sequences given the tree and the model, computed by pruning up from the tips. In 1985 he added the bootstrap: resample the sites, rebuild the tree a hundred times, and report for each branch how often it appears. For the first time a published phylogeny carried numbers that meant something about confidence.
By the end of the 1990s the Bayesian version was in place. Put a prior on trees, sample the posterior with Markov chain Monte Carlo, and read off a probability for each clade; John Huelsenbeck and Fredrik Ronquist's MrBayes made this a command line away in 2001. The machinery came straight from Bayesian statistics and Monte Carlo methods, and phylogenetics became one of their largest consumers.
Chapter IV
A Closer Look: Scoring Three Trees for Four Species
With four species there are exactly three possible unrooted trees, given by which pair is grouped: , , . Take this alignment of eight sites:
| Site | A | B | C | D | Pattern |
|---|---|---|---|---|---|
| 1 | G | G | A | A | supports |
| 2 | C | C | T | T | supports |
| 3 | T | T | G | G | supports |
| 4 | A | C | A | C | supports |
| 5 | G | T | G | T | supports |
| 6 | A | T | T | A | supports |
| 7 | C | C | C | C | constant |
| 8 | G | G | G | A | unique to D |
For a four-taxon tree, a site with two taxa in one state and two in another needs one change if the split matches the tree's internal branch, and two if it does not. A constant site needs none; a site unique to one tip needs one on every tree. So:
Parsimony picks , by one step. That margin is the whole evidence: three sites for it, two against, one for the third topology. Resample those eight sites with replacement and the winner changes often — which is exactly what Felsenstein's bootstrap measures.
Now scale up. The number of unrooted binary trees on tips is :
| Tips | Trees |
|---|---|
| 4 | 3 |
| 10 | 2,027,025 |
| 20 | |
| 50 |
Fifty species — a modest study — give more trees than there are atoms in the Earth, which is about . No search can visit them, and deciding the best one is NP-hard, so every published tree is the end point of a heuristic walk through that space.
The deeper trouble is that more sites do not always help. Under the Jukes–Cantor model, if a fraction of sites differ between two sequences, the number of substitutions per site that actually happened is
At this gives — observed difference and real change are nearly the same. At , : a quarter of the history is already hidden by sites that changed twice. At , , and at the formula diverges, because two random sequences over four bases differ at three sites in four. Past that point the sequences carry no information about time at all — they are saturated — and two long branches will look alike simply because both have gone random. That is long-branch attraction in one line, and it is why the sister group of all other animals is still disputed with whole genomes in hand.
Chapter V
One Tree, or Many?
A tree assumes that a lineage splits and the parts stay separate. Real histories break the assumption in two ways. Wayne Maddison pointed out in 1997 that even with no gene flow at all, different genes can have genuinely different trees: a polymorphism present in the ancestor gets sorted into the descendant species in whatever order chance dictates, and when speciations follow one another quickly, that order often disagrees with the species tree. Concatenating hundreds of genes does not fix this; it can produce high confidence in the wrong answer. Modern analyses model the gene trees inside the species tree, using the coalescent, rather than hoping the conflict averages out.
The second break is worse, and it is not about statistics. Genes move sideways between lineages, and in microbes they move constantly — so constantly that the tree may not be the right shape for the history at all. That is the problem of microbial phylogenomics. Where the tree does hold, calibrating it against the dated strata of palaeontology turns branching order into a chronology, and what that chronology shows over hundreds of millions of years is the subject of macroevolution.