Skip to content
Field Atlas

Atlas / Biology / The Heredity Thread

Field · Emerged 1977 – 2003

Genomics

What can be learned by reading entire genomes, and comparing them?

5 chapters4 min read5 turning points2 open problems

Branched from
Molecular Biology + Population Genetics
Branched into
Protein Structure Prediction
Figures
Frederick Sanger, Kary Mullis, Craig Venter, Francis Collins, Svante Pääbo, Jennifer Doudna, Emmanuelle Charpentier, Feng Zhang

In brief

Genomics reads the complete DNA of organisms, their genomes, and compares them. A human genome is about three billion letters long. The first took thirteen years and billions of dollars to read. Today one can be read in a day for a few hundred dollars.

Reading genomes at that scale has turned genetics from the study of single genes into the study of whole systems, and population genetics into a data science. It has also rewritten human prehistory, transformed cancer medicine, and led to tools that can edit genomes as well as read them.

Key ideas

GenomeEnters 2001 – 2003

The complete set of an organism's DNA. Only about 1–2% of the human genome codes for proteins, in roughly 20,000 genes.

DNA sequencingEnters 1977

Determining the exact order of the letters (A, C, G, T) in a stretch of DNA. Modern machines read billions of short fragments in parallel and reassemble them by computer.

PCREnters 1983 – 1985

The polymerase chain reaction copies a chosen stretch of DNA over and over, doubling it each cycle, so that a single molecule becomes billions in a few hours.

Genetic variation and GWAS

Any two people's genomes differ at a few million positions. Genome-wide association studies scan hundreds of thousands of people to link such variants to traits and diseases.

Genome editingEnters 2012

Changing a genome at a chosen position, most commonly with CRISPR–Cas9, which uses a short guide RNA to steer a DNA-cutting enzyme to a matching sequence.

Draws on other domains

Chapter I

Reading DNA

Molecular biology had shown that genes are sequences of four letters, but for twenty years the letters could hardly be read. In 1977 Frederick Sanger changed that. His method copies DNA with a small fraction of letters that stop the copying, then sorts the stopped fragments by length, and the sequence can be read off in order. Earlier that year, with a forerunner of the method, his lab had read the first complete DNA genome, a small virus of 5,386 letters.

The second tool was copying. In 1983 Kary Mullis, by his own account while driving through the California hills at night, imagined using repeated heating and cooling to double a chosen stretch of DNA again and again. His colleagues at Cetus turned the idea into a reliable method, the polymerase chain reaction, which can turn a single molecule into billions. It now underpins everything from forensic DNA to COVID tests.

Chapter II

The Human Genome

In 1990 the international Human Genome Project set out to read all three billion letters of human DNA within fifteen years. In 1998 Craig Venter's company Celera announced it would do the job faster, and privately. The race that followed, between Celera and the public consortium led by Francis Collins, ended in a truce announced at the White House in June 2000, and in rival draft papers in 2001.

The draft's great surprise was a small number: perhaps 30,000 protein-coding genes, since revised to about 20,000, not much more than a roundworm has. Complexity came from how genes are regulated and combined, not from how many there are. The project was declared complete in 2003. The hardest 8%, highly repetitive regions, was finished only in 2022.

Chapter III

Genomes and Evolution

Cheap sequencing made genomes comparable by the thousand, and population genetics became a science of whole genomes. The most startling result came from ancient DNA. In 2010 Svante Pääbo's team read a Neanderthal genome from 40,000-year-old bones and found traces of it in everyone whose ancestors lived outside Africa. Our ancestors had interbred with Neanderthals, and, a finger bone soon showed, with Denisovans. Genomes also confirmed, in exquisite detail, the tree of life that evolutionary biology had drawn from anatomy.

Chapter IV

A Closer Look: How Many Times Must a Genome Be Read?

Sequencing machines cannot read a chromosome from end to end. They read short fragments, a few hundred letters long for most of the history of genomics, from random positions, and a computer assembles the fragments by their overlaps. How much sequencing is enough?

Suppose the fragments add up to cc times the length of the genome, the coverage. Each base is then read on average cc times. Because the fragments land at random, the number of times a given base is covered follows a Poisson distribution, and the probability that it is never read at all is

P(missed)=e−c.P(\text{missed}) = e^{-c} .

Eric Lander and Michael Waterman worked out this and related formulas in 1988. They show why reading the genome once is useless:

Coverage ccFraction of bases never readFor a 3.1-billion-base genome
1×e−1≈37%e^{-1} \approx 37\%1.1 billion bases missed
3×e−3≈5%e^{-3} \approx 5\%150 million missed
8×e−8≈0.03%e^{-8} \approx 0.03\%about 1 million missed
30×e−30≈10−13e^{-30} \approx 10^{-13}none, on average

Sequencing the same total amount again gives diminishing returns, but each extra round of coverage cuts the gaps by a factor of ee. The draft human genome of 2001 used several-fold coverage and had many gaps. Clinical genome sequencing today typically uses about 30× coverage, which also allows the two copies of each chromosome to be told apart and errors to be outvoted.

Random coverage is not the only problem. About half the human genome consists of repeated sequences, and a short fragment from inside a repeat could belong to any copy of it, so the formula's gaps are the easy part. The last 8% of the genome, mostly long repeats, was finished in 2022 only with new machines that read single molecules tens of thousands of letters long. The cost of reading a human genome has meanwhile fallen from billions of dollars for the first to a few hundred dollars today.

Chapter V

Writing Genomes

In 2012 Jennifer Doudna and Emmanuelle Charpentier showed that CRISPR–Cas9, part of a bacterial defence against viruses, can be programmed to cut DNA wherever its guide RNA matches. Feng Zhang's lab and others used it in human cells within months. Genome editing became cheap and routine, and the first CRISPR-based therapy, for sickle cell disease, was approved in 2023. In 2018 He Jiankui's announcement that he had edited the genomes of twin babies was condemned almost universally, and it set off a new debate about where limits belong.

The edge of the map here is less about new techniques than about understanding what is read. Most of the heritability of common traits is still hard to pin down, and how much of the genome actually does anything is still openly disputed.

Applications

Where it is used

Open problems

Where the map runs out

Open

Missing heritability

Narrowed by very large studies but not closed, as of 2026.

Twin and family studies show that traits like height and many common diseases are highly heritable. Yet the genetic variants found by early genome-wide scans explained only a small fraction of that heritability. Where is the rest?

Why it is hard

Most of it seems to be spread across thousands of variants, each with a tiny effect, which only enormous studies can detect. Rare variants, interactions between genes, and flaws in the twin-study estimates may account for the remainder, and each is hard to measure.

What resolving it unlocks

It would determine how far disease risk can be predicted from DNA, and clarify what heritability does and does not mean.

› Sources (1)

Open

How much of the genome does anything?

Actively disputed as of 2026; estimates of the functional fraction range from under 10% to much higher.

Only about 1–2% of the human genome codes for proteins. In 2012 the ENCODE project reported biochemical activity across about 80% of the genome and called it functional. Evolutionary biologists objected strongly, since only a small fraction appears to be preserved by natural selection.

Why it is hard

"Function" means different things to a biochemist (something happens there) and to an evolutionary biologist (mutations there matter to fitness). Testing the effect of each stretch of DNA directly is slow. Much of the genome is repetitive remains of ancient viruses and mobile elements.

What resolving it unlocks

It would identify which non-coding mutations can cause disease, and settle how much of our DNA is, in effect, evolutionary debris.

› Sources (2)

Further reading

  1. Sulston, J. & Ferry, G. (2002). The Common Thread: A Story of Science, Politics, Ethics and the Human Genome. Joseph Henry Press.

    An insider's account of the public Human Genome Project and its fight to keep the data free.

  2. Pääbo, S. (2014). Neanderthal Man: In Search of Lost Genomes. Basic Books.

    The story of ancient DNA told by the scientist who pioneered it.

  3. Doudna, J. A. & Sternberg, S. H. (2017). A Crack in Creation: Gene Editing and the Unthinkable Power to Control Evolution. Houghton Mifflin Harcourt.

    CRISPR from one of its discoverers, including the ethical questions it raises.