Expanding our enterprise inference capacity with IBM Cloud and NVIDIA
Together AI
Together AI is moving enterprise AI inference onto IBM Cloud with NVIDIA B300 GPUs. It says this is the first dedicated, large-scale inference cluster of its kind there, and Together is the first customer.
Based on reporting by Together AI — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Together AI says it is working with IBM and NVIDIA to scale enterprise AI inference on IBM Cloud, starting with a large cluster of NVIDIA B300 GPUs connected by NVIDIA Spectrum-X Ethernet networking. The company calls it the first dedicated, large-scale inference cluster of its kind on IBM Cloud, and says it is the first customer using it.
The move fits Together’s bigger pitch: open models for enterprise use, not closed systems controlled by a few labs. The company says it already serves hundreds of trillions of tokens each month to more than a million developers, and that demand is still climbing. That scale is the real story here. Once usage gets that big, inference stops being a side problem and becomes the product.
Together argues that enterprises and AI-native companies keep choosing open models for two reasons: their data stays sovereign, and performance can land at a fraction of closed-model cost. The company says the economics improve as usage rises. That is the case it is making to buyers, not just to the open-source crowd.
The partnership also says something about where the market is going. Open-source AI, in Together’s telling, cannot just be clever and cheap; it has to run everywhere, fast, reliably, and with the guardrails businesses expect. This IBM Cloud cluster is one more attempt to prove that point at serious scale.
My take — AI-written commentary, not fact-checked reporting
This is the part of AI that actually matters: who can run it reliably, not who can post the flashiest demo. Open models keep winning the boring but important argument on cost, control, and deployment flexibility, which is exactly why the enterprise crowd keeps circling back. Closed labs may still own the headlines, but the plumbing is getting more interesting by the month.
Read more about this at: Together AI
Related stories
Built for Vera Rubin, NVIDIA Spectrum-6 Arrives in Gigascale AI Factories
NVIDIA · 2 months ago ·
14