TLDRocket
Sign in

Benchmarking inference at scale: coding agents

Together AI

Together released a benchmark measuring inference performance on production coding agent workloads, where Together Inference Engine delivers 31% higher throughput than TensorRT-LLM and maintains 2× better time-to-first-token at saturation of 2.5M tokens per minute. The company optimized performance through ThunderMLA (a fused kernel for Multi-head Latent Attention), custom kernels, and full-stack profiling rather than relying on single-user benchmarks. For typical coding requests of 80k-100k input tokens, Kimi K2.7 on Together costs $0.029 per request compared to $0.097 for Claude Opus, saving an engineering team of 150 people approximately $421k annually.

Why it matters

Real-world inference benchmarks for coding agents: 31% more TPS than TensorRT-LLM, 2× better TTFT at saturation, and 76% lower cost than Claude Opus 4.6.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.