TLDRocket
Sign in

The Sequence Knowledge - 928: The Missing 5%: Why Distillation Is Harder Than It Looks

TheSequence Jesus Rodriguez

A 7B model is said to keep 95% of a 70B teacher’s performance. The scary part is that the missing 5% may be the bit that actually makes it useful.

Based on reporting by TheSequence, Jesus Rodriguez — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

A new model release lands with a very neat pitch: a 7-billion-parameter student that keeps 95 percent of the performance of a 70-billion-parameter teacher. On paper, that is the sort of swap that makes people sit up. Ten times smaller. Close enough to tempt anyone building for a laptop, an agent loop, or an API margin chart.

But that 5 percent gap is doing a lot of work. It might not be a clean, even drop in quality. The student could still ace the teacher’s math-style benchmark and still fail at noticing when it is confused. It might handle familiar coding tasks, then stumble when a tool call goes sideways and needs recovery.

The problem is that distillation does not just compress answers. It may also compress judgment, and judgment is harder to measure than accuracy. A model can learn to sound right, or even to produce polished reasoning traces, without carrying over the instincts that help it cope when the task changes shape.

And that is the real warning buried in the claim. The missing 5 percent may not be a little bit of everything. It may be the bridge between good-looking outputs and a system that can actually keep going when the road disappears.

My take — AI-written commentary, not fact-checked reporting

This is why the 95 percent headline should always trigger suspicion. Open-model people love compression stories, but the last few percent are often where the useful parts hide, then quietly vanish. Fancy numbers are cheap; robust behavior is the expensive bit nobody wants to price in.

Read more about this at: TheSequence

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.