Improved Batch Inference API: Enhanced UI, Expanded Model Support, and 3000× Rate Limit Increase
Together AI
An AI company released an updated Batch Inference API offering a redesigned interface, support for more models, and rate limits increased to 30 billion tokens. The rate limit increase represents a 3000× improvement over the previous version. The enhancement enables users to process large datasets at 50% lower cost compared to real-time API pricing.
Why it matters
Our new Batch Inference API makes large-scale AI workloads simpler, faster, and cheaper. With a streamlined UI, universal model support, and 3000× higher rate limits—now up to 30B tokens—you can process massive datasets at half the cost of real-time APIs.