TLDRocket
Sign in

Kids outlearn AI—and we still don’t know why

MIT Technology Review Elise Cutts

Researchers say today’s large language models need far more training data than human children to learn language, creating a “data efficiency gap” that remains unexplained. BabyLM—the benchmark competition meant to test child-scale learning—asks teams to train models on a developmentally plausible corpus of 100 million words (10 million for the toddler track). The field is shifting toward evaluating and designing more data-efficient AI systems that better match how children acquire grammar, instead of relying only on scaling up data.

Why it matters

People have been talking to each other for at least 100,000 years, as best we can tell. And in all that time, there has been only one thing in the world that could learn a human language to perfect fluency: a human child. Now there are two. Four short years after the release of ChatGPT,…

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.