Open, convenient and predictable: Introducing Provisioned Throughput
Together AI
Together AI introduced Provisioned Throughput, a reserved inference capacity service for open-weight models with token-based pricing and a 99% uptime SLA that costs up to 90% less than Claude Opus. The service is priced at $0.05 per Provisioned Throughput Unit per minute and is currently available for MiniMax M3 and GLM-5.2 models across North America and EMEA with a one-month minimum term. Companies can now migrate production workloads from proprietary APIs to open models with guaranteed capacity and predictable pricing instead of choosing between best-effort serverless or complex dedicated inference management.
Why it matters
Provisioned Throughput gives you reserved inference capacity for frontier open models like MiniMax M3 and GLM-5.2. Token-based pricing, a 99% uptime SLA, and up to 90% lower cost than proprietary APIs. No GPU-hour math, no infrastructure to manage.