TLDRocket
Sign in

The Inference Engineering Masterclass — Philip Kiely & Ali Taha, Baseten

Latent Space

Baseten's Philip Kiely and Ali Taha discuss inference engineering as a specialized discipline for turning trained model weights into fast, reliable production APIs, covering techniques like quantization, speculative decoding, and cache-aware routing. In one GLM-5.2 experiment, quantizing more of the model increased throughput by 20% while preserving benchmark quality because errors in different layers cancelled each other out. Inference optimization has become as critical as model training itself, enabling gains of 20% to 200% and making models up to 10 times faster at scale.

Why it matters

Baseten just raised a $13B Series F and is now one of the leading kings of inference engineering. We go into everything you need to know for autoregressive and diffusion engineering.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.