Chapter I
What Is a Gene Made Of?
By 1941 genetics had shown that genes make proteins, but not what genes themselves were. Chromosomes contain both protein and DNA, and nearly everyone assumed genes were protein. DNA, with only four kinds of building block, seemed too monotonous to carry information.
Oswald Avery and his colleagues at the Rockefeller Institute showed otherwise in 1944. Extract DNA from a deadly strain of bacteria, add it to a harmless strain, and the harmless strain is permanently transformed. Destroying the proteins in the extract changes nothing. Destroying the DNA stops it. Many remained sceptical until Alfred Hershey and Martha Chase showed in 1952 that viruses inject their DNA, not their protein, into the cells they infect.
The field that took up the question was unusually cross-disciplinary. Physicists moved in, drawn by Erwin Schrödinger's 1944 book What Is Life?. Crucially, so did X-ray crystallography, the technique for reading molecular structure from how X-rays scatter, founded in physics by the Braggs in 1913.
Chapter II
The Double Helix
At King's College London, Rosalind Franklin was taking the sharpest X-ray pictures of DNA yet made, and Maurice Wilkins was working on the same problem down the corridor, with open friction between the two. In Cambridge, James Watson and Francis Crick were building models. Shown Franklin's Photo 51 by Wilkins, without her knowledge, and given her measurements through a research report, they found the structure in early 1953. Two strands wind around each other, joined by base pairs, A with T and G with C.
The structure explained itself. Each strand is a template for the other, so DNA can be copied, and in 1958 Matthew Meselson and Franklin Stahl showed that it is copied exactly that way. The sequence of bases along a strand could be a message. Franklin, who had moved on to pioneering work on viruses, died in 1958 at 37. The Nobel Prize went to the three men four years later.
Chapter III
Cracking the Code
If DNA is a message, how is it read? In 1958 Crick set out the central dogma: DNA is transcribed into RNA, RNA is translated into protein, and information never flows back out of protein. Its most important consequence is that a protein's amino-acid sequence is spelled out by its gene.
The spelling was cracked by experiment. In 1961 Marshall Nirenberg and Heinrich Matthaei made an RNA of nothing but U and got a protein of nothing but phenylalanine. Within five years every three-letter codon had been assigned. The code turned out to be nearly identical in bacteria, plants and people, powerful evidence for the common descent Darwin had argued from anatomy.
Chapter IV
A Closer Look: How Much Information Is in DNA?
DNA spells its messages in four letters, A, C, G and T. Proteins are built from twenty kinds of amino acid. How many letters does the code need per amino acid?
With one letter per amino acid, there are only 4 possible "words". With two, , still fewer than 20. With three, , enough for all twenty amino acids plus stop signals, with room to spare. George Gamow and others argued from this counting that the code must use triplets, before any experiment. Crick and Brenner's genetic experiments in 1961 confirmed it, and the spare codons explain why most amino acids have several.
Each letter carries two bits of information, since . The human genome has about 3.1 billion base pairs in one set of chromosomes, so
about the capacity of a CD-ROM. Every cell with a nucleus holds two sets, one from each parent.
It is also long. Adjacent base pairs are 0.34 nanometres apart along the helix. The two sets in a single cell, about 6.4 billion base pairs, stretch to
packed into a nucleus about 6 micrometres across, like fitting about 25 km of fine thread into a tennis ball. It is wound around protein spools and folded in loops, and still has to be unwound, copied and read in a precise order.
Copying is astonishingly accurate. The enzymes that copy DNA make roughly one error per billion letters after proofreading and repair, so a human cell division introduces only a handful of new mutations across the whole genome. That accuracy, and the rare errors it lets through, are the raw material of evolution.
Chapter V
Reading and Writing DNA
The popular version of the dogma soon needed revising. In 1970 Howard Temin and David Baltimore found that some viruses copy RNA back into DNA using reverse transcriptase, an enzyme that later made HIV understandable, and treatable.
Then biologists learned to write. In 1973 Stanley Cohen and Herbert Boyer cut DNA with enzymes that snip at specific sequences, stitched a gene into a plasmid, a small ring of bacterial DNA, and watched bacteria express it. Worried about what they had unleashed, scientists paused their own research and drew up safety rules at Asilomar in 1975. Within a decade bacteria were making human insulin. The ability to read DNA at scale, not just gene by gene, is the story of genomics.