TLDRocket
Sign in

The Sequence Knowledge #898: The Trace Is the Teacher: Distilling Reasoning Into Small Models

TheSequence Jesus Rodriguez

DeepSeek generated 800,000 worked solutions from its R1 reasoning model and fine-tuned smaller models on these step-by-step traces using standard supervised learning, achieving unexpected success without reinforcement learning or other advanced techniques. The distilled 32B model solved competition math problems significantly harder than expected for its size, while the 7B model developed emergent reasoning abilities like self-verification without explicit training. The approach challenges prior consensus that naive sequence-level imitation cannot work because student models diverge from teacher trajectories at inference time, suggesting that learning from reasoning traces may operate under different principles than simple answer imitation.

Why it matters

From the release of DeepSeek R1, distillation in reasoning models have become one of the most common techniques in frontier AI.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.