TLDRocket
Sign in

AdapTive-LeArning Speculator System (ATLAS): A New Paradigm in LLM Inference via Runtime-Learning Accelerators

Together AI

Together AI introduced ATLAS, an adaptive-learning speculator system that accelerates LLM inference by dynamically improving at runtime without manual tuning. ATLAS achieves up to 500 tokens per second on DeepSeek-V3.1 and 460 TPS on Kimi-K2, representing a 4x speedup over baseline decoding and outperforming static speculators. The system adapts continuously to changing workloads and input distributions, enabling sustained performance improvements as usage patterns evolve.

Why it matters

LLM inference that gets faster as you use it. Our runtime-learning accelerator adapts continuously to your workload, delivering 500 TPS on DeepSeek-V3.1, a 4x speedup over baseline performance without manual tuning.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.