The Sequence Knowledge #894: When the Student Started Talking Back: Distillation in the LLM Era
TheSequence Jesus Rodriguez
Knowledge distillation methods originally designed for image classification broke down when applied to language models, forcing researchers to shift from simple model compression toward capability transfer where smaller models learn to perform complex tasks with guidance from larger models. The transition occurred over approximately five years through three distinct stages that fundamentally changed how distillation operates in sequence-based tasks. This evolution reflects how language models violated the core assumptions of traditional distillation, including fixed input distributions and closed classification spaces.
Why it matters
A joiurney through the evolution of distillation for frontier models.