TLDRocket
Sign in

Together AI at NVIDIA GTC 2026: Explore our latest innovations across research and products

Together AI Covered by 2 sources

Together AI is heading to NVIDIA GTC 2026 with a pile of new integrations: Dynaimo 1.0, NemoClaw, Nemotron 3 Super, and voice AI tools. It's less one big launch than a full-stack pitch: Together wants to be the place you run NVIDIA's newest agent and reasoning tech.

Based on reporting by Together AI — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Together AI is showing up to NVIDIA's GTC conference in San Jose this March with a grab bag of announcements, and the throughline is pretty clear: everything is about making agentic AI actually deployable at production scale, not just demoable on a laptop.

The headline piece is Nemotron 3 Super, NVIDIA's new hybrid Mamba-Transformer model built for multi-agent workflows. It's a mixture-of-experts setup with 120 billion total parameters but only 12 billion active per token, paired with a million-token context window. That combination is what lets it juggle long-running reasoning tasks and multiple cooperating agents on a single GPU — useful for things like coding agents, financial analysis bots, or automated security monitoring. Together is offering it through its Dedicated Model Inference service, which is really the core of its pitch here: bring the model, we'll handle the plumbing.

On the infrastructure side, Together is folding NVIDIA's Dynamo 1.0 into its inference stack. Dynamo is NVIDIA's open-source engine for generative and agentic inference, and Together says it's already been quietly using it to squeeze more performance out of production workloads before the official 1.0 release. There's also NemoClaw, a joint project with NVIDIA that wraps up the NVIDIA OpenShell runtime — a sandboxed environment for running autonomous agents — into a one-command install. Bundle that with Together's library of over 150 optimized models, and the idea is that developers get NVIDIA's security-focused agent runtime plus Together's inference speed without stitching the two together themselves.

Voice AI gets a mention too. NVIDIA's Parakeet TDT 0.6B V3 speech recognition model is now sitting in Together's model library, aimed at developers building real-time voice agents who need low-latency transcription that doesn't fall apart under production traffic.

Beyond the product news, Together's researchers and engineering leads are doing the conference-circuit thing: sessions with Cursor and Decagon on production inference lessons, a talk from co-founder Percy Liang on open-source trust in AI research, and a booth (#1213) full of live demos and executive meetups. It's the usual GTC choreography, but the pattern worth watching is how much of Together's identity now rests on being the layer that turns NVIDIA's open models and runtimes into something a startup can actually ship.

My take — AI-written commentary, not fact-checked reporting

This reads like Together AI positioning itself as NVIDIA's favorite middleman — take the open models and runtimes NVIDIA ships, wrap them in inference infrastructure, and sell the convenience. That's a fine business, but it's also a reminder that 'open' in this ecosystem increasingly means open-weight models running on somebody's proprietary cloud stack, which is a much narrower kind of openness than the term implies. I'd rather see more scrutiny of who actually controls the inference layer than another round of GTC booth demos.

Read more about this at: Together AI

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.