TLDRocket
Sign in

Benchmarking

25 summarised stories about Benchmarking, each linking back to the original source. Browse all topics →

+ Follow this topic

Tuesday, 19 May 2026

Benchmarking inference at scale: coding agents

Together AI 3 months ago 50

Together released a benchmark measuring inference performance on production coding agent workloads, where Together Inference Engine delivers 31% higher throughput than TensorRT-LLM and maintains 2× better time-to-first-token at saturation of 2.5M tokens per minute. The company optimized performance through ThunderMLA (a fused kernel for Multi-head Latent Attention), custom kernels, and full-stack profiling rather than relying on single-user benchmarks. For typical coding requests of 80k-100k input tokens, Kimi K2.7 on Together costs $0.029 per request compared to $0.097 for Claude Opus, saving an engineering team of 150 people approximately $421k annually.

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.