Academia is for Ambition — Alex Zhang, MIT
Latent Space
MIT PhD Alex Zhang is betting on GPU kernels, agent swarms, and new ways to build AI systems. He thinks the real advantage is still human taste, not just more tokens.
Based on reporting by Latent Space — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Latent Space’s latest pick for “emerging superstar PhD” is Alex Zhang of MIT, and the draw here is not some glossy demo. It’s a string of oddly specific bets: GPU kernels, KernelBench, Recursive Language Models, and the idea that modern AI is still leaving a lot of capability on the table because the systems around the model are too primitive.
Zhang’s route into this world started with GPU Mode, the Discord community that began around learning to write GPU kernels and later broadened out. He first ran into it while interning at Snapchat and got interested in specialized kernels after hearing Tri Dao talk about FlashAttention at Princeton in 2023. From there, the path widened into competitions, lectures, and the sense that GPU programming had gone from niche obsession to something much closer to a common skill among serious AI builders.
That shift matters because Zhang doesn’t think the machine has made human expertise obsolete. On the GPU Mode leaderboard, he says many of the recent solutions are AI-generated, but the strongest systems still need people who know where the traps are. Kernels have a verification problem, reward hacking is real, and even the best results can fall apart outside the leaderboard. The same basic pattern shows up elsewhere: LLMs can help, but the people who know the problem space still steer the ship.
The bigger argument in the interview is that the future model may look simpler on the surface than it really is. Zhang talks through RLMs, context offloading, shared memory, programmatic subagent calling, persistent subagents, and harnesses as a way to generalize across tasks. His point is not that English text is going away, but that the thing you query may secretly be a swarm underneath a clean interface.
He also pushes back on the idea that brute force is the whole story. If you know how to aim the search, you can get much further with the same compute. That is a very academic kind of ambition: chase the weird bet, trust the sharp insight, and let the industry labs catch up later if they can.
My take — AI-written commentary, not fact-checked reporting
This is what academia should be doing: taking bets that look slightly ridiculous until they work. The AI field is already full of people worshipping scale, so the useful contrarian move is to ask where the system is still clumsy, brittle, or wasting its own effort. Also, “the model is really a swarm” is a very clean way to say “the interface is lying a little,” which, frankly, is most of software these days.
Read more about this at: Latent Space