TLDRocket
Sign in

The Sequence Knowledge #920: The Physics of Teaching: Distillation Scaling Laws

TheSequence Jesus Rodriguez

Distillation finally got a curve, not just folklore. Apple’s huge study says teacher size, student size, and data all obey a measurable law.

Based on reporting by TheSequence, Jesus Rodriguez — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

For years, distillation lived in the messy part of machine learning: lots of anecdotes, lots of careful experiments, and very little that looked like a rule you could plan a budget around. A stronger teacher was supposed to help. Sometimes it did. Sometimes it didn’t. More data was supposed to help too. Again, maybe. The field could explain results after the fact, but it could not say much before a long training run started burning through money.

Pretraining went through this cleanup earlier. The Kaplan scaling laws, and then Chinchilla, turned model size and data into a calculation instead of a guess. Loss started behaving like something you could fit, not something you could only debate. That mattered because it turned training into an optimization problem, and it gave the industry a number to use: roughly twenty tokens per parameter.

Distillation had been waiting for its own version of that. In early 2025, an Apple team led by Dan Busbridge ran what the source calls the most compute-intensive controlled study of distillation ever done. The setup was wide on purpose: students from 143 million to 12.6 billion parameters, teachers in a similar range, and as many as 512 billion training tokens. The point was not to show that distillation can work. Everyone already knew that. The point was to map how it works.

The result is Distillation Scaling Laws, a paper that turns a once-fuzzy practice into something much closer to physics. If student loss depends on student size and data, it also depends on the teacher, and this study finally traces that dependence as a curve. That matters because it settles the practical question practitioners kept dodging: when does a teacher help enough to justify the cost, and when is the smarter move simply to train the smaller model longer?

My take — AI-written commentary, not fact-checked reporting

This is the sort of paper that quietly humiliates a whole cottage industry of “trust the recipe” lore. Distillation has been sold like wizardry for too long, and scaling laws are the antidote: less incense, more arithmetic. Open models keep getting better when the field stops worshipping mystery and starts measuring the thing properly.

Read more about this at: TheSequence

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.