TLDRocket
Sign in

Introducing the Together AI Batch API: Process Thousands of LLM Requests at 50% Lower Cost

Together AI

Together AI launched a Batch API that processes large volumes of LLM requests asynchronously at 50% lower cost than real-time inference. The service supports up to 50,000 requests per batch file with a best-effort 24-hour completion window across 15 models including DeepSeek, Llama, and Mistral. Users can now handle non-urgent workloads like data classification and synthetic generation without consuming real-time API rate limits.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.