Why AlphaFold Didn't Solve Protein Folding — Pushmeet Kohli, Google DeepMind & Sal Candido, Biohub
At a Latent Space panel moderated by Brandon Anderson, Google DeepMind's Pushmeet Kohli and Biohub's Sal Candido argued that AlphaFold's breakthrough was a beginning rather than a solution to protein folding, and that scaling compute and data alone will not solve biology. The published conversation covers scaling laws for biological data, the tradeoff between handcrafted scientific inductive bias and general-purpose architectures, and what it would take to build predictive models of whole living systems.
Candido challenged a common assumption about scaling laws. He said the misconception is that scaling laws exist everywhere and always; much of the work is finding one in the first place. The productive situation is identifying a setting where more compute or more data reliably produces a better result, because then the effort becomes an engineering problem that can be cranked. He added that data must carry the right information and statistics for the target problem, since a model can only draw on the data it has and generalize beyond it.
Candido gave a concrete counterexample to the idea that only pristine data helps. Training a protein language model on metagenomic sequences, which he described as low quality and largely not whole proteins, still improved the model's performance at designing proteins that work and at understanding known proteins. He framed this positively but warned about the inverse pull: the temptation to simply scale up whatever data is easy to generate. He said Biohub tries to work openly with the community so that modelers and data generators together determine what data is actually needed, which is how a scaling law eventually gets found.
Kohli recast the bitter lesson away from data versus modeling. Reflecting on Rich Sutton's original lecture, which he attended while at DeepMind, he took the message as conceptual: treating yourself religiously as a modeler or as a data generation person is itself the mistake. In his framing, the problem comes first and the solution space should stay flexible, including gathering data when that is what a problem requires. He noted that early machine learning practice assumed a fixed dataset with training and test splits and just optimized the model, which he called broken if the goal is to actually solve the problem. Expertise counts as a third ingredient alongside data and modeling.
He illustrated the point with AlphaFold. The team worked with the Protein Data Bank rather than attempting to expand it by an order of magnitude, because the global investment behind that dataset was invaluable and the project lacked the expertise and resources to replicate it. That made modeling the highest-leverage investment. He contrasted this with cell genomics, where DeepMind asked what cell-by-gene data could support; after substantial work it became clear the data was not sufficient for the ambition of a virtual cell. His advice to newcomers is to think as a multidisciplinary person, understand why the problem matters, identify whether modeling, compute scaling or data generation is the binding constraint, respect constraints such as how much data can be generated or how large a model can be afforded, and fail fast.
On handcrafted architectures, Kohli defended the design work behind AlphaFold 2. Scientific intuition from biophysics and biochemistry indicated that amino acid residues influence one another, so the team baked that relationship into the model, giving it what he called an unfair advantage and making it more data-efficient because it did not have to relearn known structure. He also emphasized that curating good data is its own art: adding more data is not the same as adding good data, and duplicating existing data does not advance the state of the art. Understanding the coverage of a dataset is part of the challenge.
Candido largely agreed but added nuance about when inductive bias helps and when it hurts. With small data, more bias in the model is necessary to get results at all; as data grows, the model can find things researchers did not know, and an incorrect inductive bias can hold it back. He identified a tipping point as data and compute scale. He also pushed back on the framing of the moderator's question, saying there is substantial craft in scaling itself, including algorithmic work, not only infrastructure and inference speed. In his view the field is not post-Transformer, but architectures are being modified for specific purposes so they work better even at internet scale. As data accumulates, the recurring question is what architecture extracts the most from it at each step and scale.
On protein folding, Kohli was blunt that the problem is not solved. Science progresses by isolating a piece of a problem, and proteins are not static building blocks: they are complex, often disordered, and their shape can depend on context. He recalled discussions with John Jumper about the actual target. What AlphaFold replicated was a structure that someone had obtained and deposited in the PDB, not the true ground state of a protein or the full distribution of structures it can adopt. He said that replication happens to be useful, but it does not amount to understanding protein dynamics, and he urged continued funding for structure prediction and dynamics work.
Asked what single resource would most accelerate function, dynamics and design, Kohli pointed to cryo-EM micrographs. Drawing on his computer vision background, he reasoned that deposited structures discard information present at the source, so operating directly on the micrographs could expose dynamic and distributional detail. He said he tried this and that it requires more work, and he expects others to make progress on it.
Candido noted that these protein models are useful but solve a narrow purpose rather than the problem most people want solved, and that modeling has to move beyond individual proteins toward whole biological systems. The published discussion also touches on why protein language models may already contain scientific knowledge that has not been unlocked, why trustworthiness and uncertainty calibration matter more than full interpretability, whether frontier models could interpret other AI systems better than humans can, when AI might deliver 10x to 100x acceleration in drug discovery, and why curing all disease requires 10x breakthroughs rather than 10% improvements.
Why it matters: teams building AI for biology should treat dataset design and architecture choices as coupled decisions rather than assuming scale alone will close the gap, and should expect evaluation to shift from static pose prediction toward dynamics, disorder and whole-cell behavior. The reported performance gain from messy metagenomic sequences suggests the marginal value of data lies in coverage and information content rather than cleanliness, which affects how pipelines are built. Kohli's account implies, by my reading rather than as a stated conclusion, that benchmarks measuring agreement with deposited structures may overstate progress on biological function and design. For product teams, the practical takeaway is that interpretability may be less important than calibrated uncertainty when models are used to prioritize experiments.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
From the Bitter Lesson of AI scaling to the unsolved mysteries of protein folding, Google DeepMind’s Pushmeet Kohli and Biohub’s Sal Candido are rethinking what it takes to build AI that truly understands biology.