Chapter I
The Sequence Contains the Structure
Christian Anfinsen did the experiment that defines the problem. Ribonuclease is a small enzyme held together partly by four disulphide bridges. Treat it with urea to unfold it and a reducing agent to break the bridges, and it loses all activity; its chain is a random coil with eight free cysteines. Remove the reagents and let it sit in air, and it recovers full activity — which requires that the eight cysteines pair up in exactly the right four pairs out of 105 possible arrangements, and that the chain return to its original shape.
Nothing assisted it. No template, no energy input, no other molecule. Anfinsen drew the conclusion in 1972: the information for the structure is in the sequence, and the native structure is the thermodynamically stable one under physiological conditions. Structure prediction is therefore a well-posed problem — find the minimum of a free energy over conformations — and not, as it might have been, a question about history or machinery.
The exceptions have turned out to be informative rather than fatal. Some large proteins need chaperones to avoid aggregating on the way; prions adopt an alternative stable form and propagate it; and a third of the human proteome has regions with no fixed structure at all. But for the ordinary globular case, Anfinsen's principle holds.
Chapter II
Sixty Years of Not Solving It
The field's history between Anfinsen and 2020 is unusual in that its lack of progress was documented rigorously, by its own practitioners, every two years.
The documentation was John Moult and Krzysztof Fidelis's idea. Before 1994, a method's accuracy was reported by its author, after seeing the answer, which is not a measurement. The Critical Assessment of Structure Prediction fixed it by making the comparison blind: crystallographers release sequences whose structures they have solved but not published, groups submit predictions within weeks, and independent assessors score them once the coordinates appear. The scoring is public and so is every prediction, including the bad ones. For two decades the record showed real but slow improvement, confined largely to cases where a related structure was already known, and almost no ability to predict a fold from scratch.
Two approaches competed in that period, and the contrast between them is the point. Molecular dynamics attacks the physics directly: give every atom a position and a velocity, compute the forces from an empirical force field, integrate Newton's equations with a time step short enough to resolve a bond vibration. Martin Karplus and colleagues did it first in 1977, for 9 picoseconds of a small protein in vacuum, which was enough to establish that a protein's interior is fluid rather than rigid. Thirty years and several hardware generations later, purpose-built machines reached the millisecond and could fold small proteins from an extended chain — a genuine achievement that does not scale, for the reason the next chapter's arithmetic makes plain.
David Baker's Rosetta took the statistical route instead: assemble a candidate structure from short fragments taken from known proteins, score it with an energy function that mixes physical terms with statistics drawn from the structural database, and search. It produced the best predictions of the 2000s for proteins with no known relatives, reaching a few ångströms for small chains. It also ran backwards, which is the more surprising capability: given a target shape, search for a sequence that will adopt it. In 2003 the group designed Top7, a 93-residue protein with a fold not found in nature, and its crystal structure matched the design to 1.2 Å. Designing a protein that folds turned out to be easier than predicting how a natural one does.
Chapter III
A Closer Look: Levinthal's Numbers, and Why the Search Is Not a Search
Cyrus Levinthal made the difficulty quantitative in 1969, in two pages of a conference volume.
Take a modest protein of 100 residues. Each residue's backbone has two rotatable bonds, and allow it — generously conservatively — just 3 distinguishable conformations. The chain then has
conformations. Suppose the molecule could try one every seconds, which is about as fast as a bond rotation can occur. Searching them all takes
The universe is years old. The search would take about times longer than the universe has existed. Real proteins of this size fold in milliseconds to seconds.
So folding is not a search, and the paradox is about the shape of the energy landscape rather than about speed. If the landscape were a golf course — flat, with one hole — nothing would find the hole. The resolution, developed in the 1990s as the funnel picture, is that partly correct structures are already partly stabilised: forming a few native contacts lowers the energy, which makes the remaining search smaller. The landscape slopes towards the native state from almost everywhere, so a biased descent arrives quickly, and the chain never visits more than a tiny fraction of its conformations.
This has a consequence that is easy to miss. The funnel is a property of the sequence, and it had to be selected for. A random sequence of 100 amino acids generally does not fold at all; it aggregates or stays a coil. Evolution has produced sequences whose landscapes are funnelled, which is why the folding problem is tractable for natural proteins and why designing new ones that fold reliably is difficult.
For computation the numbers also explain what has and has not been achieved. Minimising a realistic energy function over conformations is not approachable directly, and in simplified lattice models the problem is provably NP-hard — a result from computational complexity that caps any approach based on exhaustive optimisation. Molecular dynamics attacks the physics instead, integrating every atom's motion with time steps of a femtosecond, s. Folding takes s. That is steps for one folding event of one small protein, which purpose-built hardware reached around 2010 — an achievement, and not a method for predicting structures at scale.
Chapter IV
The Answer Came From Statistics
What eventually worked used almost none of this physics. Two datasets had been accumulating. The Protein Data Bank held some 200,000 experimentally determined structures. Genome sequencing had produced hundreds of millions of protein sequences, which means that for most proteins one can assemble a deep alignment of homologues from many species.
The second dataset carries structural information in an indirect form. If two residues are in contact, a destabilising mutation at one can be compensated by a mutation at the other, so across a large family the two positions vary in a correlated way. Disentangling direct couplings from chains of indirect ones made contact prediction usable by around 2012, and contacts constrain a fold.
Rosetta had already shown what the structural database alone could support. The sequence database added the coevolution signal. What remained was a way to use both at once.
The step change came at CASP14 in 2020. AlphaFold2 reasons jointly over the sequence alignment and over a representation of every pair of residues, iterating between them, and outputs coordinates directly with a per-residue confidence estimate that turns out to be well calibrated. Its median backbone accuracy was around 1 Å, within the range that two experimental determinations of the same protein differ by. John Moult, who had run the blind assessment since 1994 and watched two decades of modest progress, said the problem was in some sense solved.
The qualification matters, and the field's own assessment is the right place to leave this. Accuracy is excellent for single folded chains with many known relatives and much weaker for sequences with no family, for complexes, for the effect of a single mutation, for the alternative conformations a protein passes through while working, and for the disordered third of the proteome that has no single structure to predict. A method trained on the endpoint also says nothing about the route: Levinthal's question, how a chain finds its way, is not answered by a system that never folds anything. What has changed is the supply. Every sequence now has a structure attached to it, which is a different science from one in which structures were produced a few hundred a year by people growing crystals.