TLDRocket
Sign in

Frontier post-training recipe review with Finbarr Timbers

Interconnects Nathan Lambert

Post-training recipes for large language models have converged on multi-teacher on-policy distillation (MOPD) as a frontier approach in 2026, replacing earlier monolithic reinforcement learning stages with domain-specialist teachers merged into a single student model. The shift occurred because single-stage RL proved expensive and created capability conflicts across math, code, and reasoning domains, while specialist models using SFT-then-RL per domain are cheaper and organizationally scalable. This architectural change, pioneered by MiMo Flash V2 in January 2026 and scaled by DeepSeek V4 and Nemotron 3 Ultra to over 10 teachers, enables labs to expand post-training complexity beyond what single-stage RL recipes like OLMo-3 could achieve.

Why it matters

"Interview" #18

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.