AI Isn’t Smarter Than a Baby—Yet
WIRED Will Knight
Researchers built a test that feeds AI the messy video a baby actually sees—and today's best models bomb it. Turns out toddlers learn faster from less data than any AI, and nobody's cracked why.
Based on reporting by WIRED, Will Knight — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Forget the benchmarks about coding or calculus for a second. The real flex in AI right now would be matching a 1-year-old, and it turns out nobody can do it. A team spanning Meta, Stanford, the University of Tokyo, and France's École Normale Supérieure just built something called the EgoBabyVLM Challenge, which force-feeds vision language models about a thousand hours of footage recorded from cameras strapped to actual infants' heads. The models choke on it.
That failure is the point. Babies don't learn from clean, labeled datasets — they learn from chaos: a parent gesturing at something off-screen, a sentence about what happened yesterday, a glance that means more than any word could. Michael Frank, a Stanford cognitive scientist who helped design the test, says the results confirm that language alone isn't enough for a model to make sense of the world the way a toddler does. There's some multimodal, tactile, physically-grounded thing happening in a developing brain that current architectures just don't replicate.
This isn't the first time cognitive scientists have used AI as a mirror to poke at old questions about the mind. Back in 2023, ETH Zurich's Ryan Cotterell built BabyLM, a challenge that had models learn grammar from roughly the same word count a 10-year-old hears — tens of millions of words instead of the trillions modern LLMs devour. Surprisingly, transformer models did fine, which is awkward for anyone still clinging to Noam Chomsky's idea that syntax is hardwired into human brains. But language turned out to be the easy part. MIT's Joshua Tenenbaum points out that BabyLM never produced anything resembling common sense about physics, social behavior, or what other people are thinking. Transformers are pattern-matching machines, he says, and pattern-matching alone doesn't get you the kind of reasoning a toddler picks up by age two.
Some early experiments hint at a path forward. In 2024, a basic vision language model trained purely on one infant's head-camera footage figured out what a ball is. And this year Frank's team tested a model specifically built to track causality and how objects move and interact over time, again using baby-cam data — it learned physical dynamics far better than standard architectures. The implication is that baking in the right biases, the kind evolution seemingly built into the human brain, might matter more than throwing an ocean of data at a generic transformer.
Brendan Lake at Princeton, who worked on the infant-video experiments, calls EgoBabyVLM
My take — AI-written commentary, not fact-checked reporting
The AI industry loves to brag about scale — more chips, more tokens, more billions burned — while a toddler with zero GPUs quietly outlearns it on efficiency alone. If that doesn't humble the
Read more about this at: WIRED