TLDRocket
Sign in

Benchmarking inference at scale: coding agents

Together AI

Together released a benchmark measuring inference performance on production coding agent workloads, where Together Inference Engine delivers 31% higher throughput than TensorRT-LLM and maintains 2× better time-to-first-token at saturation of 2.5M tokens per minute. The company optimized performance through ThunderMLA (a fused kernel for Multi-head Latent Attention), custom kernels, and full-stack profiling rather than relying on single-user benchmarks. For typical coding requests of 80k-100k input tokens, Kimi K2.7 on Together costs $0.029 per request compared to $0.097 for Claude Opus, saving an engineering team of 150 people approximately $421k annually.

Why it matters

Real-world inference benchmarks for coding agents: 31% more TPS than TensorRT-LLM, 2× better TTFT at saturation, and 76% lower cost than Claude Opus 4.6.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.