The Inference Engineering Masterclass — Philip Kiely & Ali Taha, Baseten
Latent Space
Baseten's Philip Kiely and Ali Taha discuss inference engineering as a specialized discipline for turning trained model weights into fast, reliable production APIs, covering techniques like quantization, speculative decoding, and cache-aware routing. In one GLM-5.2 experiment, quantizing more of the model increased throughput by 20% while preserving benchmark quality because errors in different layers cancelled each other out. Inference optimization has become as critical as model training itself, enabling gains of 20% to 200% and making models up to 10 times faster at scale.
Why it matters
Baseten just raised a $13B Series F and is now one of the leading kings of inference engineering. We go into everything you need to know for autoregressive and diffusion engineering.