Chapter I
Reading DNA
Molecular biology had shown that genes are sequences of four letters, but for twenty years the letters could hardly be read. In 1977 Frederick Sanger changed that. His method copies DNA with a small fraction of letters that stop the copying, then sorts the stopped fragments by length, and the sequence can be read off in order. Earlier that year, with a forerunner of the method, his lab had read the first complete DNA genome, a small virus of 5,386 letters.
The second tool was copying. In 1983 Kary Mullis, by his own account while driving through the California hills at night, imagined using repeated heating and cooling to double a chosen stretch of DNA again and again. His colleagues at Cetus turned the idea into a reliable method, the polymerase chain reaction, which can turn a single molecule into billions. It now underpins everything from forensic DNA to COVID tests.
Chapter II
The Human Genome
In 1990 the international Human Genome Project set out to read all three billion letters of human DNA within fifteen years. In 1998 Craig Venter's company Celera announced it would do the job faster, and privately. The race that followed, between Celera and the public consortium led by Francis Collins, ended in a truce announced at the White House in June 2000, and in rival draft papers in 2001.
The draft's great surprise was a small number: perhaps 30,000 protein-coding genes, since revised to about 20,000, not much more than a roundworm has. Complexity came from how genes are regulated and combined, not from how many there are. The project was declared complete in 2003. The hardest 8%, highly repetitive regions, was finished only in 2022.
Chapter III
Genomes and Evolution
Cheap sequencing made genomes comparable by the thousand, and population genetics became a science of whole genomes. The most startling result came from ancient DNA. In 2010 Svante Pääbo's team read a Neanderthal genome from 40,000-year-old bones and found traces of it in everyone whose ancestors lived outside Africa. Our ancestors had interbred with Neanderthals, and, a finger bone soon showed, with Denisovans. Genomes also confirmed, in exquisite detail, the tree of life that evolutionary biology had drawn from anatomy.
Chapter IV
A Closer Look: How Many Times Must a Genome Be Read?
Sequencing machines cannot read a chromosome from end to end. They read short fragments, a few hundred letters long for most of the history of genomics, from random positions, and a computer assembles the fragments by their overlaps. How much sequencing is enough?
Suppose the fragments add up to times the length of the genome, the coverage. Each base is then read on average times. Because the fragments land at random, the number of times a given base is covered follows a Poisson distribution, and the probability that it is never read at all is
Eric Lander and Michael Waterman worked out this and related formulas in 1988. They show why reading the genome once is useless:
| Coverage | Fraction of bases never read | For a 3.1-billion-base genome |
|---|---|---|
| 1× | 1.1 billion bases missed | |
| 3× | 150 million missed | |
| 8× | about 1 million missed | |
| 30× | none, on average |
Sequencing the same total amount again gives diminishing returns, but each extra round of coverage cuts the gaps by a factor of . The draft human genome of 2001 used several-fold coverage and had many gaps. Clinical genome sequencing today typically uses about 30× coverage, which also allows the two copies of each chromosome to be told apart and errors to be outvoted.
Random coverage is not the only problem. About half the human genome consists of repeated sequences, and a short fragment from inside a repeat could belong to any copy of it, so the formula's gaps are the easy part. The last 8% of the genome, mostly long repeats, was finished in 2022 only with new machines that read single molecules tens of thousands of letters long. The cost of reading a human genome has meanwhile fallen from billions of dollars for the first to a few hundred dollars today.
Chapter V
Writing Genomes
In 2012 Jennifer Doudna and Emmanuelle Charpentier showed that CRISPR–Cas9, part of a bacterial defence against viruses, can be programmed to cut DNA wherever its guide RNA matches. Feng Zhang's lab and others used it in human cells within months. Genome editing became cheap and routine, and the first CRISPR-based therapy, for sickle cell disease, was approved in 2023. In 2018 He Jiankui's announcement that he had edited the genomes of twin babies was condemned almost universally, and it set off a new debate about where limits belong.
The edge of the map here is less about new techniques than about understanding what is read. Most of the heritability of common traits is still hard to pin down, and how much of the genome actually does anything is still openly disputed.