Kids outlearn AI—and we still don’t know why
MIT Technology Review Elise Cutts
Researchers say today’s large language models need far more training data than human children to learn language, creating a “data efficiency gap” that remains unexplained. BabyLM—the benchmark competition meant to test child-scale learning—asks teams to train models on a developmentally plausible corpus of 100 million words (10 million for the toddler track). The field is shifting toward evaluating and designing more data-efficient AI systems that better match how children acquire grammar, instead of relying only on scaling up data.
Why it matters
People have been talking to each other for at least 100,000 years, as best we can tell. And in all that time, there has been only one thing in the world that could learn a human language to perfect fluency: a human child. Now there are two. Four short years after the release of ChatGPT,…