Making Knowledge Distillation Cheap Enough to Run at Scale
Hugging Face
Efficient Knowledge Distillation for LLMs proposes caching teacher top-K logits offline and using a fused, chunked KL-divergence loss to avoid building full vocabulary-by-sequence probability grids during training. The work reports cutting peak VRAM from about 250GB with dense KL to about 128GB with fused chunked KL, enabling long-context distillation on a single H200-class GPU. As a result, training cost drops enough for more large-scale experiments, and the authors claim near-lossless match to online distillation at 8K context while scaling to 32K–256K contexts with much lower memory.