Skip to content
Field Atlas

Atlas / Biology / The Molecular Structure Thread

Field · Emerged 1965 – 2017

Structural Biology

Once a molecule's atoms can be located, what does the arrangement explain — and how large an assembly can be solved?

4 chapters6 min read6 turning points1 open problem

Branched from
Protein Crystallography + Molecular Biology
Branched into
Membrane Biophysics + Molecular Machines + Protein Structure Prediction
Figures
David Phillips, Kurt Wüthrich, Hartmut Michel, Johann Deisenhofer, Robert Huber, Thomas Steitz, Peter Moore, Venkatraman Ramakrishnan, Ada Yonath, Jacques Dubochet, Joachim Frank, Richard Henderson

In brief

Myoglobin showed what a protein looks like. The next question was whether a structure explains anything, and the answer arrived in 1965 with lysozyme, the first enzyme to be solved. Its active site is a groove with a sugar chain lying in it and two acidic side chains positioned on either side of the bond that gets broken. The mechanism was readable from the picture, and the field's central claim — that function follows from shape, in detail — was established.

What followed was a sixty-year climb in the size of what could be solved. A membrane protein, long thought impossible to crystallise, was done in 1985. Nuclear magnetic resonance made it possible to determine structures in solution rather than in crystals. In 2000 three groups solved the ribosome, a machine of some quarter of a million atoms, and found that its catalytic centre contains no protein at all within 18 Å — the ribosome is a ribozyme, which settled a long argument about what came first. Then electron microscopy, after thirty years as the method of last resort, acquired detectors good enough to work from single particles, and structures of objects that will not crystallise became routine.

Key ideas

Structure–function relationEnters 1965 – 1967

The claim that what a molecule does is determined by, and legible from, the arrangement of its atoms. Lysozyme's groove with two catalytic residues flanking the scissile bond was the first case where a mechanism was read off a map.

Active siteEnters 1965 – 1967

A pocket or cleft, usually a small fraction of the molecule's volume, in which a few precisely positioned residues do the chemistry. The rest of the protein is largely scaffolding that holds them in place.

Structural databaseEnters 1971

A public archive of coordinates, deposited on publication. It converted structures from individual results into a dataset that could be mined for folds, motifs and statistics, and it is the training set every prediction method depends on.

Solution NMREnters 1982 – 1986

Determining a structure from nuclear magnetic resonance measurements of distances between atoms in solution, with no crystal. It works for smaller proteins, reports on flexibility rather than averaging it away, and is the main source of information about disordered regions.

Single-particle reconstructionEnters 2012 – 2017

Build a three-dimensional map by averaging tens of thousands of low-dose electron micrographs of individual copies of a molecule, each frozen in a random orientation. No crystal is needed, and different conformations can be separated rather than merged.

RibozymeEnters 2000 – 2001

A catalyst made of RNA. The ribosome's peptide-bond-forming site is one, which means protein synthesis is carried out by RNA and is evidence that RNA catalysis preceded protein catalysis.

Draws on other domains

Chapter I

A Mechanism Read Off a Map

The first protein structure, myoglobin, explained very little. Myoglobin stores oxygen, and seeing where the haem sits confirmed that it has a pocket for it, which was not news. The question was whether structures would ever explain chemistry.

Lysozyme answered it. David Phillips's group solved the enzyme at 2 Å in 1965, and then did the decisive experiment: solved it again with a sugar inhibitor bound. The substrate lies in a long groove across the molecule's face. Two acidic residues sit on either side of the bond that gets cut — Glu35 positioned to donate a proton to the leaving oxygen, Asp52 positioned to stabilise the positive charge that develops on the sugar. And the sugar in the fourth subsite cannot fit without being distorted out of its relaxed shape, towards the flattened geometry it must adopt in the transition state.

Everything in that paragraph is visible in the map. An enzyme accelerates a reaction by binding the transition state better than the substrate, and here was a picture of it being done. After lysozyme, "solve the structure" became the standard move for understanding any protein, and the rest of the field is about extending the range of what can be solved.

Chapter II

Three Barriers, Removed in Turn

Crystals. Kurt Wüthrich showed in the 1980s that a structure can be obtained in solution instead. Nuclear magnetic resonance can report which protons are within about 5 Å of each other; collect enough such constraints, assign each resonance to a specific atom in the sequence, and compute the conformations consistent with the list. The method is limited to smaller proteins, and it has an advantage crystallography lacks: it sees motion, and reports a family of conformations rather than one.

Membranes. A membrane protein has a hydrophobic belt that must be in contact with lipid, so it will not dissolve in water, and the detergents that keep it soluble interfere with crystal packing. The consensus was that these proteins could not be crystallised, which was awkward, since they are about a quarter of the proteome and most drug targets. Hartmut Michel found conditions that worked for a bacterial photosynthetic reaction centre, and Johann Deisenhofer and Robert Huber solved it in 1985. The structure shows the pigments down which an electron hops after a photon is absorbed, at spacings that explain the direction and speed of the transfer — a result that belongs as much to wave optics and quantum mechanics as to biology.

Size. By the late 1990s the frontier was the ribosome: two subunits, three RNA chains of several thousand bases, more than fifty proteins, a quarter of a million atoms, and a crystal that diffracts badly. Three groups solved it between 1999 and 2001, in a race described frankly by one of the participants.

Chapter III

A Closer Look: Eighteen Ångströms of Nothing

The ribosome structures settled a question that biochemistry had been unable to touch. The peptidyl transferase centre is where an incoming amino acid is joined to the growing chain — the chemical step at the heart of all protein synthesis. Was it catalysed by one of the ribosome's fifty-odd proteins, or by its RNA?

The argument from the structure is a distance measurement. In the 2.4 Å map of the large subunit, with a transition-state analogue bound in the active site, the nearest atom of any protein side chain is about 18 Å from the site of the reaction. For comparison:

DistanceWhat is at that range
1.5 Åa covalent bond
2.8 Åa hydrogen bond
3–4 Åvan der Waals contact
18 Åfive water molecules' worth of nothing

No chemistry reaches 18 Å. Acid–base catalysis requires a proton donor in hydrogen-bonding range; electrostatic stabilisation falls off with distance and, through water, is screened within a few ångströms. Whatever the ribosomal proteins do — and they do stabilise the structure, and assist assembly — they cannot be performing the catalysis. The catalytic site is lined by RNA bases, and the ribosome is therefore a ribozyme.

The consequence reaches back four billion years. Protein synthesis cannot have required proteins to begin with, so the machine that makes proteins is made of the other polymer. Crick, Orgel and Woese had each suggested in the 1960s that RNA came first, on the grounds that it can both carry information and fold; catalytic RNAs were found in the 1980s, which showed it was possible; the ribosome structure showed that the most fundamental process in the cell still works that way.

There is a second lesson in the numbers, about what resolution buys. At 5 Å the ribosome is a shape, and the argument above cannot be made. At 2.4 Å individual bases, bound waters and the analogue's geometry are all placed, and an 18 Å gap becomes a measurement rather than an impression. The difference between the two maps is a factor of four in the number of reflections measured — and about fifteen years of work.

Chapter IV

Pictures Instead of Crystals

The last barrier fell for a mundane reason: better cameras. Electron microscopy of frozen biological molecules had been possible since Jacques Dubochet worked out in the early 1980s how to freeze a sample so fast that the water becomes glass rather than ice, and Joachim Frank had developed the mathematics for taking tens of thousands of images of individual particles lying in random orientations and averaging them into a three-dimensional map. The resulting maps were blurry, and the field was nicknamed blobology.

What changed between 2012 and 2014 was the detector. Direct electron detectors record individual electrons and read out fast enough to split an exposure into frames, so the drift of the specimen during exposure — previously an irreducible blur — can be tracked and corrected. Resolutions crossed 3 Å, and then went further.

The consequences have been larger than an improvement in resolution. No crystal is needed, so complexes that had resisted crystallisation for decades were solved within months, including the spliceosome and many membrane receptors. And because each image is of one particle, a sample containing several conformations can be sorted computationally into separate maps instead of being averaged into an uninterpretable mean — so a machine can be caught in several states of its cycle, which is what molecular machines requires.

What structural biology still cannot handle is the third of the proteome that has no fixed structure at all. Disordered regions are conserved, functional, and invisible to every method in this chapter, and the question of what should replace a list of coordinates for them is open.

Applications

Where it is used

  • Drug discovery

    Designing against a pocket

    Structure-based design starts from the shape and chemistry of a target's binding site and builds a molecule to fit it. HIV protease inhibitors were produced this way within a decade of the enzyme's structure; so were the influenza neuraminidase inhibitors, and the kinase inhibitors that dominate modern oncology. The approach does not remove the need for medicinal chemistry, and it changes where the chemistry starts.

    › Sources (2)
    • Wlodawer, A. & Vondrasek, J. (1998). Inhibitors of HIV-1 protease. Annual Review of Biophysics 27: 249–284.
    • Blundell, T. L. (2017). Protein crystallography and drug discovery. IUCrJ 4: 308–321.
  • Antibiotics

    Where the ribosome-binding drugs bind

    Over half of clinically used antibiotics act on the bacterial ribosome, and before 2000 nobody knew where. The structures showed the binding sites for the macrolides, tetracyclines, aminoglycosides and oxazolidinones, the differences from the human ribosome that give selectivity, and the mutations by which resistance arises — which turned resistance from an observation into something addressable by design.

    › Sources (1)
    • Wilson, D. N. (2014). Ribosome-targeting antibiotics and mechanisms of bacterial resistance. Nature Reviews Microbiology 12: 35–48.
  • Origin of life

    Evidence for an RNA world

    That the ribosome's catalytic centre is pure RNA is among the strongest pieces of evidence that RNA catalysis preceded protein catalysis, since the machine that makes all proteins cannot itself have required proteins to work. It is a structural result bearing on a question about events four billion years ago, and it is cited in every account of life's origin.

    › Sources (2)
    • Cech, T. R. (2000). The ribosome is a ribozyme. Science 289: 878–879.
    • Fox, G. E. (2010). Origin and evolution of the ribosome. Cold Spring Harbor Perspectives in Biology 2: a003483.

Open problems

Where the map runs out

Open

What to do about proteins with no fixed structure

Open as of 2026; no agreed representation exists for an ensemble rather than a structure.

A third or more of human proteins contain long regions with no stable fold, and many are disordered throughout. They are not defective: disorder is conserved, and these regions mediate a large share of regulatory interactions and form the condensates that organise parts of the cell. Crystallography sees nothing of them, and the question of what should replace a set of coordinates — an ensemble of what size, weighted how, described by which parameters — has no settled answer.

Why it is hard

Experiments report averages over an ensemble, and many different ensembles give the same average, so the inverse problem is badly underdetermined. Simulations can generate ensembles but depend on force fields calibrated on folded proteins, and there is no benchmark of known answers to validate against, because the quantity to be predicted is not a single structure.

What resolving it unlocks

Regulation, signalling and the formation of membraneless compartments all involve these regions, and most are undruggable by methods that assume a pocket. A workable description of a disordered ensemble would open a third of the proteome to structural reasoning.

› Sources (2)
  • van der Lee, R. et al. (2014). Classification of intrinsically disordered regions and proteins. Chemical Reviews 114: 6589–6631.
  • Bonomi, M., Heller, G. T., Camilloni, C. & Vendruscolo, M. (2017). Principles of protein structural ensemble determination. Current Opinion in Structural Biology 42: 106–116.

Further reading

  1. Branden, C. & Tooze, J. (1999). Introduction to Protein Structure, 2nd edition. Garland.

    The standard illustrated introduction to folds and what they do.

  2. Ramakrishnan, V. (2018). Gene Machine. Basic Books.

    A first-hand account of the race to solve the ribosome, unusually candid about competition.

  3. Kühlbrandt, W. (2014). The resolution revolution. Science 343: 1443–1444.

    Two pages on why cryo-EM suddenly worked, written as it was happening.