TLDRocket
Sign in

Understanding the brain with AI-driven explanations and experiments

Microsoft Jianfeng Gao

Researchers built a way to turn AI brain-prediction models into plain-English theories, then test them in an fMRI scanner. Turns out the AI's guesses about what your brain regions actually do can be checked and confirmed, not just trusted blindly.

Based on reporting by Microsoft, Jianfeng Gao — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

For a decade now, LLMs have been eerily good at predicting how your brain lights up when you listen to a story. Feed one the same text a person hears in a scanner, and it can match the activity of specific cortex patches with real precision. The problem: nobody can explain why. The model is just millions of parameters, and it tells you a region responds to language without saying whether that region cares about food, places, numbers, or something nobody's thought to ask about yet.

A team from Microsoft Research, UC Berkeley, UCSF, and Columbia decided to make the black box talk. Their method, generative causal testing, works in two moves. First, an LLM looks at the phrases that most strongly drive a brain region's predicted response and boils them down into a short label, something like "food preparation" or "location names." Second, and this is the clever part, another LLM writes brand-new stories built specifically to trigger that explanation, and three human subjects go back into the scanner to hear them. If the target region actually lights up more than it does for ordinary text, the explanation survives contact with reality instead of just sounding plausible.

The validation held up across all three subjects, and the explanations were most reliable exactly where the underlying prediction models were strongest, which is a nice sanity check in itself. But the real payoff came from applying GCT to open questions. Three neighboring place-related regions, RSC, PPA, and OPA, have long been lumped together as functionally similar. GCT pulled them apart by generating stories designed to activate one while suppressing the others, and found that RSC specifically cares about proper-noun locations like Tokyo or Connecticut, not location in general.

Even more striking, the method turned up prefrontal micro-regions nobody had mapped before: one tuned to dialogue markers like "said" or "told," another to clock times such as "one o'clock," and a third to measurements like "50 feet." Nobody was hunting for a clock-time detector in the prefrontal cortex. It surfaced because the method let researchers propose a hypothesis and test it immediately, rather than waiting for a lucky guess.

The authors frame this as bigger than neuroscience. Any field drowning in predictive-but-opaque models, and there are plenty, faces the same fork: accept the black box, or find a way to distill it into something a human can read and then check. GCT is a decent template for the second option, and it's the kind of unglamorous, methodological work that tends to matter more in five years than the splashier headlines do.

My take — AI-written commentary, not fact-checked reporting

This is the good, boring kind of AI progress that never trends on social media but actually earns its keep: using an LLM not to replace scientific reasoning but to generate falsifiable hypotheses a human can then go check with real data. I'd take ten papers like this over another benchmark leaderboard flex any day, and I hope funding bodies notice that interpretability work is where the actual scientific value of these models gets unlocked.

Read more about this at: Microsoft

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.