Biohub and partners commit $1.8B to virtual-cell data
The Neuron ● Covered by 3 sources
Biohub and partners are putting $1.8B into data for a “virtual cell.” The bet is that better measurements, not bigger models alone, will make biology predictable.
Based on reporting by The Neuron — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Biohub is widening its Virtual Biology Initiative into a much bigger push, with Google DeepMind, Meta, Isomorphic Labs and U.S. agencies now in the mix. The goal is bluntly ambitious: build the data needed for a “virtual cell,” an AI model that can forecast how living cells will react to changes.
That starts with a lot of wet-lab work. Biohub launched the five-year effort in April with a $500 million pledge, including $400 million for its own data generation and measurement tools, plus $100 million for a broader international effort. DeepMind, Meta and Isomorphic Labs are adding $300 million together. The Department of Energy is putting in more than $500 million over five years, and NIH is taking on the job of coordinating datasets and research resources created through more than $500 million in earlier federal spending.
The $1.8 billion headline is not just fresh cash. It also folds in existing data, compute and measurement technology, along with earlier commitments and previously funded resources. Biohub is also bringing in the Allen Institute, Broad Institute, Wellcome Sanger Institute, Human Cell Atlas and Human Protein Atlas, while NVIDIA will provide computing infrastructure, software and technical help. The plan is to make those pieces behave like one system through shared standards, common identifiers and a single access point.
The reason this is harder than protein folding is simple: a cell is not a fixed object. AlphaFold showed what AI can do for molecular structure. A virtual cell has to deal with what happens when biology changes. That means measurements across many conditions, from advanced imaging to molecular, cellular and tissue engineering, so models can capture biology’s molecular, spatial and dynamic sides.
The DOE side adds supercomputing, advanced experimental facilities and automated labs, which should help with both data generation and model testing. There’s a commercial angle too. Axios reported that commercial partners get one year of exclusive access to the data they produce before it goes public, while Biohub’s Alex Rives told Reuters that embargoes can help pull in private money. Government-funded work in parallel won’t have that restriction.
The real test is still ahead. Rives expects a first dataset in about a year and accurate predictive models within five years, according to Reuters. But a Nature Methods study of 27 methods across 29 datasets found generalization is still a problem, especially in unfamiliar cellular settings. The dream is nice; the lab will be rude about it.
My take — AI-written commentary, not fact-checked reporting
This is the right bet, and it’s overdue. Biology won’t be cracked by chanting “more parameters” at it; it needs instruments, standards and a pile of expensive reality. The awkward truth is that the next big AI win in medicine may come from less model worship and more boring infrastructure.
Read more about this at: The Neuron