In plain English
Proteins are the molecular machines of life. Haemoglobin carries oxygen. Insulin regulates blood sugar. Collagen holds your skin together. Antibodies neutralise pathogens. Enzymes catalyse every chemical reaction in your metabolism. Everything your body does, it does with proteins.
A protein is built from amino acids, small molecules that come in 20 varieties, strung together in a chain. The sequence of amino acids in the chain is encoded in DNA. That part, reading the sequence from the gene, was understood and made routine by the 1970s. The hard part came next.
A chain of amino acids is not a protein. It is a floppy string. To become functional, it must fold, contorting itself into a precise three-dimensional shape in which some parts form tight helices, others flat sheets, others irregular loops, and the whole thing settles into a configuration so specific that a single misfolded protein can cause disease. This folding happens spontaneously, in microseconds, driven by chemistry. Every protein folds into the same shape every time, reliably, without instruction.
The problem: given only the sequence of amino acids, predict the three-dimensional structure the protein will fold into.
This was Anfinsen's hypothesis (Nobel Prize, 1972): structure is determined by sequence alone. The information to fold is entirely contained in the chain. But the number of possible conformations a chain of even 100 amino acids could adopt is astronomical. Levinthal calculated in 1969 that if a protein tried every possible configuration randomly, it would take longer than the age of the universe to find the right one. Yet proteins fold in microseconds. They do not try randomly. They follow a folding pathway. But what pathway?
This became one of the central unsolved problems of biology, with implications for understanding every disease that involves protein malfunction, which is most of them.
Five things to file under "wait, what?"
-
AlphaFold 2 solved it in 2020, and the scientific community's reaction was disbelief. The Critical Assessment of Protein Structure Prediction (CASP) competition has been running since 1994, providing blind tests where teams predict structures of proteins whose experimental structure is known but not released. In CASP14 in 2020, AlphaFold 2 achieved a median score of 92.4 GDT (out of 100), the equivalent of the experimentally determined structure by most measures. The previous best score, from the 2018 competition where AlphaFold 1 had already shocked the field, was around 45. Demis Hassabis, CEO of DeepMind, received the Nobel Prize in Chemistry in 2024. An AI trained on the known structures in the Protein Data Bank solved a problem that had resisted fifty years of structural biology.
-
Misfolded proteins cause some of the most devastating diseases known. Alzheimer's disease involves the misfolding and aggregation of amyloid-beta and tau proteins into plaques and tangles in the brain. Parkinson's disease involves misfolded alpha-synuclein. Creutzfeldt-Jakob disease (and BSE, "mad cow disease") are caused by prions, misfolded proteins that propagate by causing other copies of the same protein to misfold. Type 2 diabetes involves the aggregation of a misfolded pancreatic protein called IAPP. Understanding how these proteins misfold, and finding molecules that prevent it, depends on knowing their structure. AlphaFold has produced predicted structures for over 200 million proteins, covering every known protein from sequenced organisms, and released them all for free.
-
The Protein Data Bank contained around 100,000 experimentally determined protein structures when AlphaFold launched. It now has access to over 200 million AlphaFold predictions. Experimental structure determination, using X-ray crystallography, cryo-electron microscopy, or NMR, takes months to years per protein and requires specialist equipment. AlphaFold predicts a structure in minutes on a laptop. The Database of AlphaFold predictions, maintained by EMBL-EBI, is freely downloadable and freely searchable. Structural biology shifted from a specialised experimental discipline to something closer to a database query.
-
The human genome encodes approximately 20,000 proteins, of which fewer than half had known structures before AlphaFold. Many of the remaining proteins could not be crystallised for X-ray diffraction, or were too large or flexible for other methods. AlphaFold provided predicted structures for all of them. This does not replace experimental validation. AlphaFold predictions are not perfect, particularly for disordered regions or for how proteins change shape when they bind other molecules, but they provide a starting point that has accelerated drug discovery and basic research.
-
A second revolution arrived in 2024: RoseTTAFold All-Atom can now predict how proteins interact with other molecules. Knowing a protein's shape is only part of the drug discovery problem. You also need to know how it will bind to a small molecule drug, or how it will interact with another protein, or how a mutation changes the binding site. RoseTTAFold All-Atom, from David Baker's lab at the University of Washington (Baker shared the 2024 Nobel with Hassabis), can predict protein interactions with DNA, RNA, small molecules, and other proteins. It is a step toward predicting the full molecular dynamics of the cell.
The full story
Why it was so hard
The protein folding problem sounds straightforward: given the sequence, find the structure. The difficulty is the scale of the search space.
A typical protein has a few hundred amino acids. Each amino acid in a chain has some freedom to rotate relative to its neighbours, a small number of degrees of freedom per residue. The total number of possible conformations is combinatorially explosive: even with generous simplifications, the number of possible shapes a 100-amino-acid protein could adopt is estimated at 10^47. The age of the universe in nanoseconds is around 10^26.
This is Levinthal's paradox: proteins cannot fold by random search, yet they do fold. The resolution is that protein folding is not a random search. It follows a funnel-shaped energy landscape. Most configurations have high free energy (are unstable). Perturbations that reduce free energy are thermodynamically favoured. The protein does not try random shapes; it rolls downhill on an energy landscape toward the minimum.
The difficulty for computational prediction was that the energy functions governing this landscape are complex products of the physical chemistry of all the atoms involved: electrostatics, van der Waals interactions, hydrogen bonds, hydrophobic effects, entropy. Calculating them accurately enough to predict the global minimum was, for fifty years, beyond available methods.
What AlphaFold does
AlphaFold 2 does not solve the physics from first principles. It learns from data. The key insight was the use of multiple sequence alignments: taking a target protein and finding all the homologous proteins (similar sequences from other organisms) that evolution has preserved. Positions in the sequence that vary across species tend to be structurally tolerant; positions that are conserved tend to be structurally critical. Pairs of positions that co-evolve often do so because they are in physical contact in the structure.
AlphaFold 2 uses a transformer architecture, similar to the models underlying large language models, to process these evolutionary signals together with the raw sequence, learning representations of the relationship between sequence and structure from the ~170,000 experimentally determined structures in the Protein Data Bank. The result is a model that generalises far beyond its training data, predicting structures for proteins with no close homologue at accuracy levels that compete with experiment.
The fact that it worked at the level it did surprised even its creators. DeepMind's CEO Demis Hassabis, announcing the result, described it as a potential solution to a problem he had defined as central to his scientific agenda for decades.
What comes after structure
Knowing structure is necessary but not sufficient for understanding biology or designing drugs. Proteins are dynamic: they flex, breathe, and change shape when they bind partners or are modified. A static predicted structure captures one conformation, but the biologically relevant question is often about conformational change.
The next generation of tools is addressing this. Molecular dynamics simulations, guided by AlphaFold structures, can sample the conformational ensemble of a protein over time. AlphaFold 3 (2024) extends predictions to protein-DNA, protein-RNA, and protein-small-molecule complexes. OpenFold, ESMFold (Meta), and other open-source models continue to extend and refine the approach.
Drug discovery, the identification of small molecules that bind to and modulate protein function, has traditionally been a slow, expensive, failure-prone process. The bottleneck of unknown protein structures is substantially reduced. This does not make drug discovery easy: binding prediction, off-target effects, toxicity, pharmacokinetics, and clinical translation remain hard. But the structural foundation on which all of this sits has been transformed.
The Nobel
In October 2024, David Baker, Demis Hassabis, and John Jumper received the Nobel Prize in Chemistry. Baker received recognition for his decades of work on protein design, creating novel proteins with specified functions. Hassabis and Jumper received recognition for AlphaFold. It was a rare case of a Nobel Prize awarded for work that the entire relevant scientific community had agreed, without serious dissent, was revolutionary.
The CASP community, which had been running blind tests for twenty-six years, watching the gradual progress of computational methods, described AlphaFold 2's performance as "a solution to the protein structure prediction problem." That framing, from the people who had spent careers working on it, captures the scale of what happened.
Go deeper
- The Protein Folding Problem β Nature 2020 review β the AlphaFold 2 paper itself, open access
- AlphaFold database β freely searchable predicted structures for over 200 million proteins
- Life's Blueprint by Berton RouechΓ© β for background on proteins and molecular biology
- The Gene: An Intimate History by Siddhartha Mukherjee β broader context on the molecular biology of life
- How AlphaFold solved protein folding β Veritasium β YouTube
- DeepMind AlphaFold explained β YouTube
- What is protein folding? β TED-Ed β YouTube