TLDRocket
Sign in

Why AlphaFold Didn't Solve Protein Folding — Pushmeet Kohli, Google DeepMind & Sal Candido, Biohub

Latent Space

AlphaFold was a breakthrough, but DeepMind and Biohub say protein folding still isn’t solved. The real lesson: biology needs the right data and the right problem, not just bigger models.

Based on reporting by Latent Space — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

AlphaFold made a huge splash, but Pushmeet Kohli and Sal Candido are arguing that it was never the end of the story. On a Latent Space panel, the DeepMind and Biohub researchers kept circling back to the same point: in biology, scale helps, but scale without the right data, the right architecture, and a clear problem is mostly expensive optimism.

Candido’s view is blunt. There’s no magic law that appears just because a dataset gets bigger. First you have to find a setting where extra compute and extra data actually improve the result. And if the data is missing the information you need, the model won’t conjure it out of nowhere. He even pointed to protein language models trained on metagenomic sequences that aren’t pristine — or, in many cases, even complete proteins — yet still improve design and understanding of real proteins.

Kohli pushed the same idea from another angle. The mistake, in his telling, is treating yourself as either a model person or a data person. Start with the problem, then decide what the problem demands. For AlphaFold, DeepMind leaned on existing datasets because it did not have the resources to massively expand the Protein Data Bank. For cell genomics, though, Kohli said the team eventually concluded that the data was not ready for the larger ambition of a virtual cell.

That leads to a quieter but more interesting argument: biology still needs hand-built judgment. Kohli said AlphaFold benefited from scientific intuition from biophysics and biochemistry, especially the idea that amino acid residues influence one another rather than acting alone. Candido agreed, but added that there is craft in scaling too. Bigger models, more compute, and more data are not automatic wins; sometimes the architecture has to change as well, and not every problem belongs in a standard transformer-shaped box.

Both men also drew a hard line under the “protein folding solved” narrative. AlphaFold helped with static structure prediction, but protein dynamics and disorder remain open, and richer data like cryo-EM micrographs may be needed to move beyond single-protein snapshots. The deeper bet here is that useful biology will come from combining modeling, data generation, and scientific expertise — not worshipping any one of them.

My take — AI-written commentary, not fact-checked reporting

This is the sensible anti-hype take, which means it will annoy people who want a single hero model and a victory lap. Biology has been battered enough by “just scale it” thinking already; the field needs fewer poster slogans and more people willing to admit when the data is wrong, thin, or simply absent. Open, messy collaboration beats model cosplay every time.

Read more about this at: Latent Space

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.