TLDRocket
Sign in

Frontier post-training recipe review with Finbarr Timbers

Interconnects Nathan Lambert

Post-training recipes for large language models have converged on multi-teacher on-policy distillation (MOPD) as a frontier approach in 2026, replacing earlier monolithic reinforcement learning stages with domain-specialist teachers merged into a single student model. The shift occurred because single-stage RL proved expensive and created capability conflicts across math, code, and reasoning domains, while specialist models using SFT-then-RL per domain are cheaper and organizationally scalable. This architectural change, pioneered by MiMo Flash V2 in January 2026 and scaled by DeepSeek V4 and Nemotron 3 Ultra to over 10 teachers, enables labs to expand post-training complexity beyond what single-stage RL recipes like OLMo-3 could achieve.

Why it matters

"Interview" #18

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.