TLDRocket
Sign in

Introducing the Together AI Batch API: Process Thousands of LLM Requests at 50% Lower Cost

Together AI

Together AI launched a Batch API that processes large volumes of LLM requests asynchronously at 50% lower cost than real-time inference. The service supports up to 50,000 requests per batch file with a best-effort 24-hour completion window across 15 models including DeepSeek, Llama, and Mistral. Users can now handle non-urgent workloads like data classification and synthetic generation without consuming real-time API rate limits.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.