AdapTive-LeArning Speculator System (ATLAS): A New Paradigm in LLM Inference via Runtime-Learning Accelerators
Together AI
Together AI introduced ATLAS, an adaptive-learning speculator system that accelerates LLM inference by dynamically improving at runtime without manual tuning. ATLAS achieves up to 500 tokens per second on DeepSeek-V3.1 and 460 TPS on Kimi-K2, representing a 4x speedup over baseline decoding and outperforming static speculators. The system adapts continuously to changing workloads and input distributions, enabling sustained performance improvements as usage patterns evolve.
Why it matters
LLM inference that gets faster as you use it. Our runtime-learning accelerator adapts continuously to your workload, delivering 500 TPS on DeepSeek-V3.1, a 4x speedup over baseline performance without manual tuning.