TLDRocket
Sign in

Open, convenient and predictable: Introducing Provisioned Throughput

Together AI

Together AI launched Provisioned Throughput, reserved inference capacity for open models like MiniMax M3 and GLM-5.2 with a 99% uptime SLA. It's the middle ground companies wanted: predictable pricing and guarantees without managing GPUs themselves.

Based on reporting by Together AI — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Together AI just closed a gap that's been bugging enterprise buyers of open-weight models for a while now. Until this week, if you wanted to run MiniMax M3 or GLM-5.2 in production, you had two choices: serverless inference, which is cheap and easy but best-effort, or dedicated inference on your own GPUs, which gives you control but dumps GPU-hour math and capacity planning on your team. Provisioned Throughput is the thing in between — you buy Provisioned Throughput Units, each a fixed slice of guaranteed token-per-minute capacity, priced at $0.05 per PTU per minute, backed by a 99% uptime SLA.

The pitch is basically: give us the same kind of token-based capacity deal you already have with your closed-model provider, but point it at frontier open models instead. No infrastructure to configure, same API as their serverless and dedicated offerings, just reserved capacity you can count on. It's live today for MiniMax M3 and GLM-5.2, with availability in North America, EMEA, and other regions, and a one-month minimum commitment. Together says more models are coming.

The economics are the real hook here. Input tokens, cached input tokens, and output tokens each burn down a PTU at different rates — on MiniMax M3, one PTU handles 138,840 input tokens per minute, 694,200 cached input tokens per minute, or 23,140 output tokens per minute, in whatever mix your traffic actually needs. At full utilization, Together works that out to roughly $0.36 per million input tokens and $2.16 per million output tokens on M3, against $5 and $25 list for Claude Opus 4.8 — up to 90% cheaper, by their numbers.

What's driving this isn't abstract. Together says token volume through its APIs has gone from 30 billion to more than 400 trillion a month in nine months, and a meaningful chunk of that used to run on closed APIs. Companies making the switch report 6-20x lower inference costs. MiniMax's Linda Sheng framed it as giving teams a way to shift volume off closed APIs with commitments they can actually plan a business around, which is a fairly candid admission that the old serverless-versus-dedicated split just wasn't cutting it for production workloads.

Together is positioning this as the missing rung on the ladder: serverless to prototype, Provisioned Throughput once you need guarantees at scale, dedicated inference if you need a custom or fine-tuned model with full control. Whether it actually delivers 99% uptime under real traffic is something customers will find out over the next few months, not something a launch post can prove on its own.

My take — AI-written commentary, not fact-checked reporting

This is the unglamorous but necessary plumbing that decides whether the open-model shift is real or just a cost experiment on the side. If companies can get closed-API-style guarantees on models that cost a fraction as much, the incentive to keep paying Anthropic or OpenAI list price for routine agent work basically evaporates. The bigger tell here isn't the SLA, it's that token volume reportedly went from 30 billion to over 400 trillion a month in nine months — that's the number worth watching, not the marketing copy around it.

Read more about this at: Together AI

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.