Skip to content
Field Atlas

Atlas / Biology / The Molecular Structure Thread

Field · Emerged 1961 – 2021

Protein Structure Prediction

If a protein's shape is determined by its sequence, can the shape be computed from the sequence — and if so, how?

4 chapters7 min read6 turning points1 open problem

Branched from
Structural Biology + Genomics
Branched into
Not yet surveyed past here
Figures
Christian Anfinsen, Cyrus Levinthal, Martin Karplus, J. Andrew McCammon, John Moult, Krzysztof Fidelis, David Baker, John Jumper, Demis Hassabis

In brief

Christian Anfinsen showed in the early 1960s that a protein unfolded in the test tube refolds, unaided, into exactly the shape it had before. The sequence therefore contains the structure, and the structure is presumably the lowest-energy state. That turned prediction into a well-posed computational problem, and it stayed unsolved for sixty years.

The obstacle is the size of the search. Cyrus Levinthal pointed out in 1969 that a chain of a hundred residues has astronomically many conformations, so a protein cannot find its native state by trying them, and a computer certainly cannot. Nature's answer is that folding is funnelled rather than searched. The computational answer, when it arrived, came from a different direction entirely: not from physics but from statistics over the 200,000 structures that crystallographers had deposited, and the far larger number of sequences that genome projects had produced. Related sequences from different species carry, in their patterns of correlated mutation, information about which residues touch each other. AlphaFold2 learned to read that information in 2020, and reached accuracies comparable to experiment — for folded single chains, which is not the whole problem.

Key ideas

Thermodynamic hypothesisEnters 1961 – 1973

The native structure is the one of lowest free energy for the sequence under physiological conditions, reachable without external information. Anfinsen's refolding experiments are the evidence, and the exceptions — chaperone-dependent proteins, prions, disordered regions — are instructive.

Levinthal's paradoxEnters 1969

A random search through a protein's possible conformations would take longer than the age of the universe, yet folding takes milliseconds. The native state is therefore found by a biased descent, not by enumeration.

Folding funnelEnters 1969

The energy landscape is not flat with one hole in it but shaped like a funnel: partly correct structures are already partly stabilised, so the search is downhill almost everywhere. This is a property the sequence must have been selected for.

Coevolution signalEnters 2020 – 2022

If two residues are in contact, a mutation at one is often compensated by a mutation at the other, so their variation across homologous sequences is correlated. A deep enough family of related sequences therefore encodes a map of which residues touch.

Blind assessmentEnters 1994

Predictions are submitted for proteins whose structures have been determined but not released, so the comparison cannot be tuned after the fact. Running this every two years since 1994 is what made claims of progress in the field checkable.

Molecular dynamicsEnters 1977 – 2010

Integrating Newton's equations for every atom of a protein and its surrounding water, with an empirical force field. It simulates the folding process rather than predicting the endpoint, and is limited by the gap between femtosecond time steps and millisecond folding times.

Draws on other domains

Chapter I

The Sequence Contains the Structure

Christian Anfinsen did the experiment that defines the problem. Ribonuclease is a small enzyme held together partly by four disulphide bridges. Treat it with urea to unfold it and a reducing agent to break the bridges, and it loses all activity; its chain is a random coil with eight free cysteines. Remove the reagents and let it sit in air, and it recovers full activity — which requires that the eight cysteines pair up in exactly the right four pairs out of 105 possible arrangements, and that the chain return to its original shape.

Nothing assisted it. No template, no energy input, no other molecule. Anfinsen drew the conclusion in 1972: the information for the structure is in the sequence, and the native structure is the thermodynamically stable one under physiological conditions. Structure prediction is therefore a well-posed problem — find the minimum of a free energy over conformations — and not, as it might have been, a question about history or machinery.

The exceptions have turned out to be informative rather than fatal. Some large proteins need chaperones to avoid aggregating on the way; prions adopt an alternative stable form and propagate it; and a third of the human proteome has regions with no fixed structure at all. But for the ordinary globular case, Anfinsen's principle holds.

Chapter II

Sixty Years of Not Solving It

The field's history between Anfinsen and 2020 is unusual in that its lack of progress was documented rigorously, by its own practitioners, every two years.

The documentation was John Moult and Krzysztof Fidelis's idea. Before 1994, a method's accuracy was reported by its author, after seeing the answer, which is not a measurement. The Critical Assessment of Structure Prediction fixed it by making the comparison blind: crystallographers release sequences whose structures they have solved but not published, groups submit predictions within weeks, and independent assessors score them once the coordinates appear. The scoring is public and so is every prediction, including the bad ones. For two decades the record showed real but slow improvement, confined largely to cases where a related structure was already known, and almost no ability to predict a fold from scratch.

Two approaches competed in that period, and the contrast between them is the point. Molecular dynamics attacks the physics directly: give every atom a position and a velocity, compute the forces from an empirical force field, integrate Newton's equations with a time step short enough to resolve a bond vibration. Martin Karplus and colleagues did it first in 1977, for 9 picoseconds of a small protein in vacuum, which was enough to establish that a protein's interior is fluid rather than rigid. Thirty years and several hardware generations later, purpose-built machines reached the millisecond and could fold small proteins from an extended chain — a genuine achievement that does not scale, for the reason the next chapter's arithmetic makes plain.

David Baker's Rosetta took the statistical route instead: assemble a candidate structure from short fragments taken from known proteins, score it with an energy function that mixes physical terms with statistics drawn from the structural database, and search. It produced the best predictions of the 2000s for proteins with no known relatives, reaching a few ångströms for small chains. It also ran backwards, which is the more surprising capability: given a target shape, search for a sequence that will adopt it. In 2003 the group designed Top7, a 93-residue protein with a fold not found in nature, and its crystal structure matched the design to 1.2 Å. Designing a protein that folds turned out to be easier than predicting how a natural one does.

Chapter IV

The Answer Came From Statistics

What eventually worked used almost none of this physics. Two datasets had been accumulating. The Protein Data Bank held some 200,000 experimentally determined structures. Genome sequencing had produced hundreds of millions of protein sequences, which means that for most proteins one can assemble a deep alignment of homologues from many species.

The second dataset carries structural information in an indirect form. If two residues are in contact, a destabilising mutation at one can be compensated by a mutation at the other, so across a large family the two positions vary in a correlated way. Disentangling direct couplings from chains of indirect ones made contact prediction usable by around 2012, and contacts constrain a fold.

Rosetta had already shown what the structural database alone could support. The sequence database added the coevolution signal. What remained was a way to use both at once.

The step change came at CASP14 in 2020. AlphaFold2 reasons jointly over the sequence alignment and over a representation of every pair of residues, iterating between them, and outputs coordinates directly with a per-residue confidence estimate that turns out to be well calibrated. Its median backbone accuracy was around 1 Å, within the range that two experimental determinations of the same protein differ by. John Moult, who had run the blind assessment since 1994 and watched two decades of modest progress, said the problem was in some sense solved.

The qualification matters, and the field's own assessment is the right place to leave this. Accuracy is excellent for single folded chains with many known relatives and much weaker for sequences with no family, for complexes, for the effect of a single mutation, for the alternative conformations a protein passes through while working, and for the disordered third of the proteome that has no single structure to predict. A method trained on the endpoint also says nothing about the route: Levinthal's question, how a chain finds its way, is not answered by a system that never folds anything. What has changed is the supply. Every sequence now has a structure attached to it, which is a different science from one in which structures were produced a few hundred a year by people growing crystals.

Applications

Where it is used

  • Genomics

    A structure for every sequence

    Genome projects produce sequences far faster than crystallography produces structures, and before 2021 the great majority of proteins in any organism had no structural information at all. Predicted structures now cover most known sequences, which has let whole proteomes be annotated by fold, remote relationships be detected through shape where sequence similarity had vanished, and previously unassignable proteins from environmental samples be given candidate functions.

    › Sources (2)
    • Varadi, M. et al. (2022). AlphaFold Protein Structure Database. Nucleic Acids Research 50: D439–D444.
    • van Kempen, M. et al. (2024). Fast and accurate protein structure search with Foldseek. Nature Biotechnology 42: 243–246.
  • Complexity theory↗ Mathematics · Computational Complexity

    Folding is hard, in the technical sense

    Even in drastically simplified models — a chain of hydrophobic and polar beads on a two-dimensional lattice — finding the minimum-energy conformation is NP-hard. The result does not say that real proteins cannot fold, since nature is not searching for a global optimum in a worst-case instance; it says that no algorithm can be relied on to find the optimum for every sequence, which is why methods that work are statistical rather than exhaustive.

    › Sources (2)
    • Berger, B. & Leighton, T. (1998). Protein folding in the hydrophobic-hydrophilic (HP) model is NP-complete. Journal of Computational Biology 5: 27–40.
    • Unger, R. & Moult, J. (1993). Finding the lowest free energy conformation of a protein is an NP-hard problem. Bulletin of Mathematical Biology 55: 1183–1198.
  • Protein engineering

    Designing molecules that do not exist

    Running prediction backwards — searching for a sequence that will adopt a specified shape — has produced proteins with folds not found in nature, enzymes catalysing reactions with no natural counterpart, self-assembling cages used as vaccine scaffolds, and binders made to order against chosen targets. The 2024 Nobel Prize in Chemistry recognised both the design and the prediction halves of this work.

    › Sources (2)
    • Kuhlman, B. et al. (2003). Design of a novel globular protein fold with atomic-level accuracy. Science 302: 1364–1368.
    • Watson, J. L. et al. (2023). De novo design of protein structure and function with RFdiffusion. Nature 620: 1089–1100.

Open problems

Where the map runs out

Open

Predicting the states a protein moves between

Open as of 2026; accurate single-structure prediction has not extended to ensembles or to mutational effects.

Proteins work by changing shape: a channel opens, a receptor flips between active and inactive, a motor cycles, an enzyme closes over its substrate. Prediction methods return one structure, usually the most populated, and do not reliably say what the alternatives are, how much they cost, or which ligand shifts the balance. The related failure is quantitative: a method that places every atom correctly can still mispredict whether a single amino acid substitution destabilises the protein.

Why it is hard

The training data are crystal structures, which are biased towards whichever state crystallised, and the free-energy differences between functional states are a few kilocalories per mole — below the resolution of the statistical signal that makes single-structure prediction work. Physics-based simulation can in principle supply these differences and is limited by force-field accuracy and by the gap between simulated and biological timescales.

What resolving it unlocks

Drug design needs the state a molecule binds and the energetic consequence of a mutation; genetics needs to know which of the millions of observed human variants matter. Both are questions about differences between states rather than about a single structure.

› Sources (2)
  • Lane, T. J. (2023). Protein structure prediction has reached the single-structure frontier. Nature Methods 20: 170–173.
  • Chakravarty, D. & Porter, L. L. (2022). AlphaFold2 fails to predict protein fold switching. Protein Science 31: e4353.

Further reading

  1. Dill, K. A. & MacCallum, J. L. (2012). The protein-folding problem, 50 years on. Science 338: 1042–1046.

    A clear statement of what the problem is, written before it was substantially solved.

  2. Jumper, J. et al. (2021). Highly accurate protein structure prediction with AlphaFold. Nature 596: 583–589.

    The paper itself; unusually readable about why the architecture is shaped as it is.

  3. Moore, P. B. et al. (2022). The protein-folding problem: not yet solved. Science 375: 507.

    One page on what remains, from four structural biologists.