TLDRocket
Sign in

Announcing General Availability of Together Instant Clusters, offering ready to use, self-service NVIDIA GPUs

Together AI

Together AI just made renting big GPU clusters as easy as clicking a button, no procurement hell required. That matters because AI labs waste days wiring up hardware instead of actually training models.

Based on reporting by Together AI — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Together AI is taking the aim-and-fire approach to GPU infrastructure with the general release of Instant Clusters, a self-service system that lets companies spin up anywhere from a single 8-GPU node to hundreds of interconnected GPUs in minutes rather than weeks. The pitch is simple: stop treating tightly networked GPU clusters like some exotic procurement project involving tickets, contracts, and manual setup, and start treating them like the rest of cloud computing — API-first, predictable, and fast.

The technical package is genuinely loaded. Clusters ship pre-wired with NVIDIA Quantum-2 InfiniBand for scale-out training and NVLink/NVLink Switch inside each node, plus a stack of infrastructure plumbing — GPU Operator, ingress controllers, Cert Manager, NVIDIA Network Operator — that teams usually burn days assembling themselves. Orchestration comes via Kubernetes or Slurm, with SSH access when needed, and everything can be provisioned through console, CLI, API, or tools like Terraform and SkyPilot for multi-cloud setups. Together says every node gets burn-in testing and NVLink checks before a job even starts, with continuous monitoring afterward to catch a bad NIC or overheating GPU before it quietly wrecks a training run.

Pricing is refreshingly plain: hourly rates from $1.76 to $5.50 per GPU-hour depending on hardware (NVIDIA HGX H100 up through the newer B200) and commitment length, free data transfer, and $0.16 per GiB-month for shared storage. No hidden fees, no long-term lock-in unless you want the discount for a 1-week-to-3-month term.

The use cases Together highlights are telling. Fractal AI Research Lab describes bursty workloads — spin up a big cluster for 24 to 48 hours, hammer through a training job, then scale back down. Latent Health is running large-scale reinforcement learning on clinical question sets to distill smaller models that reportedly beat bigger foundation models on specific tasks. And Together's own chief scientist, Tri Dao, the guy behind FlashAttention, makes the underlying argument explicit: raw FLOPs stopped being the bottleneck a while ago. The real constraint is how fast you can get a clean, well-networked cluster online so researchers spend their time on architecture and data instead of babysitting infrastructure.

My take — AI-written commentary, not fact-checked reporting

This is Together AI correctly betting that infrastructure friction, not GPU scarcity, is the next competitive battleground — and honestly it's overdue, because H100s sitting idle behind a two-week procurement queue is a worse problem than most people admit. I'd watch whether the reliability claims hold up under real multi-week training runs rather than demo bursts, since that's where clusters historically fall apart. Still, treating GPU clusters like disposable cloud infrastructure instead of precious pets is the right instinct, and it's the kind of boring plumbing work that actually moves the field forward more than another benchmark headline.

Read more about this at: Together AI

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.