TLDRocket
Sign in

🔬Causal Models Need Causal Data - Xaira’s X-Cell model for Drug Discovery (Bo Wang & Ci Chu, Chief Discovery Officer & Chief AI Scientist)

Latent Space RJ Honicky

Xaira built a huge new gene-editing dataset called X-Atlas because their AI model hit a wall no amount of extra compute could fix. Turns out predicting what drugs do to cells needs way richer data, not just bigger models.

Based on reporting by Latent Space, RJ Honicky — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

There's a moment in AI research that shows up as a flat line on a chart, and it's more damning than any error message. Bo Wang and Ci Chu at Xaira Therapeutics hit exactly that wall: their model's test loss stopped improving past 1.5 billion parameters, even as training loss kept dropping. Scale up to 3.1 billion parameters and the model actually falls off the trend line. That's the tell. No amount of extra compute or extra parameters was going to fix it, because the problem wasn't the model. It was the data.

The data in question came from CELLxGENE, a massive public resource built by the Chan Zuckerberg Institute that catalogs RNA expression across 168 million cells, tracking somewhere between 20,000 and 30,000 genes per cell. It's a roughly four-trillion-entry matrix, and it has done for virtual cell modeling what the Protein Data Bank did for structural biology. Bo Wang's own scGPT, one of the most cited models trained on this kind of data, became a foundation stone for a lot of what followed. But there's a catch: CELLxGENE tells you what cell states look like, not what causes them. Gene expression changes are tangled together, correlated in ways that make it nearly impossible to untangle cause from effect just by observing.

So Chu and Wang's teams built something new instead of trying to squeeze more out of what existed. X-Atlas is a dataset generated through CRISPR-based perturbation experiments run in parallel, millions of times over, each one nudging a single gene and recording what happens upstream and downstream. That's the causal signal missing from observational data like CELLxGENE. Feed a model enough of these targeted interventions and, in theory, it can start predicting what a drug or gene edit will actually do to a cell, rather than just describing what cells tend to look like. X-Cell is the model trained on top of that, and unlike its predecessor, it scales the way you'd want a model to scale: more parameters and more compute actually buy you more performance.

What's striking is the shape of the investment behind this. Nobody's disclosed exact figures, but data collection and lab infrastructure likely ran into the tens of millions of dollars, while compute, headcount, and modeling research probably cost a few million on top of that. That ratio is backwards from typical AI pretraining economics, where compute dominates. It looks more like an RL rollout budget than a data-hungry language model budget, which says something about where the actual scarcity sits in biology right now. It's not FLOPs. It's wet-lab experiments that generate information no public database currently contains.

The promotions that followed — Chu to Chief Discovery Officer, Wang to Chief AI Scientist — happened after this work was already underway, which tells you how seriously Xaira is treating the causal-data bet as core strategy rather than a side project. Whether X-Cell generalizes cleanly to real human cells in real lab settings, and whether it can meaningfully beat the stubbornly strong linear baselines that have dogged this field, remains the open question the team is still chasing.

My take — AI-written commentary, not fact-checked reporting

This is the most sensible thing I've seen come out of the virtual-cell hype cycle: instead of throwing more parameters at correlational data and hoping causality falls out, Xaira went and generated causal data on purpose. It's a good reminder that in biology, unlike in text, you can't just scrape your way to ground truth — somebody has to run the CRISPR experiment. I'd bet this data-first approach quietly outperforms a dozen flashier foundation-model announcements this year, precisely because it's less flashy.

Read more about this at: Latent Space

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.