Skip to content
Field Atlas

Atlas / Biology / The Heredity Thread

Field · Emerged 1944 – 1966

Molecular Biology

What are genes made of, and how do they build a living cell?

5 chapters4 min read6 turning points1 open problem

Branched from
Genetics + Biochemistry
Branched into
Genomics + Structural Biology
Figures
Oswald Avery, Rosalind Franklin, Maurice Wilkins, James Watson, Francis Crick, Marshall Nirenberg, Howard Temin, David Baltimore, Stanley Cohen, Herbert Boyer

In brief

Molecular biology explains heredity and the workings of cells in terms of molecules. Genes are stretches of DNA, a double helix whose two strands carry the same information in complementary form. The cell copies DNA into RNA, and reads RNA three letters at a time to build proteins, the molecular machines that do almost everything else.

In about twenty years, from 1944 to 1966, it answered what a gene is, how it is copied and how it is read. By the 1970s scientists could cut, copy and move genes between organisms, the beginning of biotechnology.

Key ideas

The double helixEnters 1953

DNA is two strands wound around each other, held together by pairs of bases: A with T, G with C. Each strand determines the other, which is how DNA can be copied.

The central dogmaEnters 1958

Crick's summary: information flows from DNA to RNA to protein, and not back from protein. Retroviruses later showed that it can flow from RNA back to DNA, which Crick's own statement had allowed but most biologists had not expected.

The genetic codeEnters 1961 – 1966

The table that maps each three-letter RNA "codon" to one of twenty amino acids, or to "stop". It is nearly the same in every organism on Earth.

Transcription and translationEnters 1958

Transcription copies a gene from DNA into messenger RNA. Translation, on the ribosome, reads that RNA to assemble a protein.

Recombinant DNAEnters 1973

DNA cut from one source and joined to another, for example a human gene inserted into a bacterium, which then makes the human protein.

Draws on other domains

Chapter I

What Is a Gene Made Of?

By 1941 genetics had shown that genes make proteins, but not what genes themselves were. Chromosomes contain both protein and DNA, and nearly everyone assumed genes were protein. DNA, with only four kinds of building block, seemed too monotonous to carry information.

Oswald Avery and his colleagues at the Rockefeller Institute showed otherwise in 1944. Extract DNA from a deadly strain of bacteria, add it to a harmless strain, and the harmless strain is permanently transformed. Destroying the proteins in the extract changes nothing. Destroying the DNA stops it. Many remained sceptical until Alfred Hershey and Martha Chase showed in 1952 that viruses inject their DNA, not their protein, into the cells they infect.

The field that took up the question was unusually cross-disciplinary. Physicists moved in, drawn by Erwin Schrödinger's 1944 book What Is Life?. Crucially, so did X-ray crystallography, the technique for reading molecular structure from how X-rays scatter, founded in physics by the Braggs in 1913.

Chapter II

The Double Helix

At King's College London, Rosalind Franklin was taking the sharpest X-ray pictures of DNA yet made, and Maurice Wilkins was working on the same problem down the corridor, with open friction between the two. In Cambridge, James Watson and Francis Crick were building models. Shown Franklin's Photo 51 by Wilkins, without her knowledge, and given her measurements through a research report, they found the structure in early 1953. Two strands wind around each other, joined by base pairs, A with T and G with C.

The structure explained itself. Each strand is a template for the other, so DNA can be copied, and in 1958 Matthew Meselson and Franklin Stahl showed that it is copied exactly that way. The sequence of bases along a strand could be a message. Franklin, who had moved on to pioneering work on viruses, died in 1958 at 37. The Nobel Prize went to the three men four years later.

Chapter III

Cracking the Code

If DNA is a message, how is it read? In 1958 Crick set out the central dogma: DNA is transcribed into RNA, RNA is translated into protein, and information never flows back out of protein. Its most important consequence is that a protein's amino-acid sequence is spelled out by its gene.

The spelling was cracked by experiment. In 1961 Marshall Nirenberg and Heinrich Matthaei made an RNA of nothing but U and got a protein of nothing but phenylalanine. Within five years every three-letter codon had been assigned. The code turned out to be nearly identical in bacteria, plants and people, powerful evidence for the common descent Darwin had argued from anatomy.

Chapter IV

A Closer Look: How Much Information Is in DNA?

DNA spells its messages in four letters, A, C, G and T. Proteins are built from twenty kinds of amino acid. How many letters does the code need per amino acid?

With one letter per amino acid, there are only 4 possible "words". With two, 42=164^2 = 16, still fewer than 20. With three, 43=644^3 = 64, enough for all twenty amino acids plus stop signals, with room to spare. George Gamow and others argued from this counting that the code must use triplets, before any experiment. Crick and Brenner's genetic experiments in 1961 confirmed it, and the spare codons explain why most amino acids have several.

Each letter carries two bits of information, since 4=224 = 2^2. The human genome has about 3.1 billion base pairs in one set of chromosomes, so

3.1×109×2 bits=6.2×109 bits≈775 megabytes,3.1 \times 10^9 \times 2 \text{ bits} = 6.2 \times 10^9 \text{ bits} \approx 775 \text{ megabytes} ,

about the capacity of a CD-ROM. Every cell with a nucleus holds two sets, one from each parent.

It is also long. Adjacent base pairs are 0.34 nanometres apart along the helix. The two sets in a single cell, about 6.4 billion base pairs, stretch to

6.4×109×0.34×10−9 m≈2.2 m,6.4 \times 10^9 \times 0.34 \times 10^{-9} \text{ m} \approx 2.2 \text{ m} ,

packed into a nucleus about 6 micrometres across, like fitting about 25 km of fine thread into a tennis ball. It is wound around protein spools and folded in loops, and still has to be unwound, copied and read in a precise order.

Copying is astonishingly accurate. The enzymes that copy DNA make roughly one error per billion letters after proofreading and repair, so a human cell division introduces only a handful of new mutations across the whole genome. That accuracy, and the rare errors it lets through, are the raw material of evolution.

Chapter V

Reading and Writing DNA

The popular version of the dogma soon needed revising. In 1970 Howard Temin and David Baltimore found that some viruses copy RNA back into DNA using reverse transcriptase, an enzyme that later made HIV understandable, and treatable.

Then biologists learned to write. In 1973 Stanley Cohen and Herbert Boyer cut DNA with enzymes that snip at specific sequences, stitched a gene into a plasmid, a small ring of bacterial DNA, and watched bacteria express it. Worried about what they had unleashed, scientists paused their own research and drew up safety rules at Asilomar in 1975. Within a decade bacteria were making human insulin. The ability to read DNA at scale, not just gene by gene, is the story of genomics.

Applications

Where it is used

Open problems

Where the map runs out

Open

The origin of the genetic code

Open as of 2026.

Why does life use this particular table of 64 codons to 20 amino acids, nearly the same in every organism? Is the code a "frozen accident" of early history, or did chemistry or selection shape it?

Why it is hard

The code was fixed before the last common ancestor of all life, more than 3.5 billion years ago, and no organisms from before then survive. Statistical patterns in the code, such as its resistance to errors, suggest some optimisation but cannot establish how it arose.

What resolving it unlocks

It would illuminate the transition from chemistry to biology, and it bears on how the origin of life itself happened.

› Sources (1)

Recently resolved

The protein folding problem

Structure prediction largely solved by AlphaFold 2 (2020–21). How proteins physically fold, and how they misfold, remain open.

A protein is a chain of amino acids that folds itself into a precise three-dimensional shape, and its shape determines its job. For fifty years, since Christian Anfinsen showed that the sequence alone determines the fold, predicting the shape from the sequence was one of biology's grand challenges.

Why it is hard

A chain can take an astronomically large number of shapes, yet it finds its fold in milliseconds. Physics-based simulation was far too slow. The breakthrough came from machine learning trained on all the structures determined experimentally, which predicts shapes without fully explaining how folding happens.

What resolving it unlocks

Predicted structures for nearly every known protein now speed up drug design and the study of disease. Understanding folding itself, and misfolding in diseases like Alzheimer's, is still open.

› Sources (1)

Further reading

  1. Judson, H. F. (1979). The Eighth Day of Creation: Makers of the Revolution in Biology. Simon & Schuster.

    The definitive history of molecular biology's founding years, built on interviews with its makers.

  2. Maddox, B. (2002). Rosalind Franklin: The Dark Lady of DNA. HarperCollins.

    A biography that restores Franklin's role in the double helix.

  3. Alberts, B. et al. (2022). Molecular Biology of the Cell (7th ed.). W. W. Norton.

    The standard textbook, comprehensive and well illustrated.