TLDRocket
Sign in

Together AI Delivers Top Speeds for DeepSeek-R1-0528 Inference on NVIDIA Blackwell

Together AI

Together AI launched NVIDIA Blackwell GPU support for its inference platform, achieving 334 tokens/second throughput for DeepSeek-R1-0528 models as of July 17, 2025. The company's optimized inference stack combines custom GPU kernels, a proprietary inference engine, speculative decoding, and lossless quantization to deliver a 32 tokens/second speed improvement over baseline deployments and 2.3x to 2.8x faster performance than prior-generation H200 GPUs. Customers can now access this optimized inference capability through Together's serverless endpoints and dedicated endpoints, with the latter offering up to 84 tokens/second additional speedup for production workloads.

Why it matters

Together AI inference is now among the world’s fastest, most capable platforms for running open-source reasoning models like DeepSeek-R1 at scale, thanks to our new inference engine designed for NVIDIA HGX B200.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.