TLDRocket
Sign in

Nvidia launches a smaller, faster Nemotron model and a router to put it to work

The New Stack Frederic Lardinois Covered by 4 sources

Nvidia just shipped a smaller Nemotron model and a router to pick which AI model does the work. It’s pushing open models for speed and lower costs, not just benchmark bragging rights.

Based on reporting by The New Stack, Frederic Lardinois — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Nvidia on Tuesday added Nemotron 3.5 Lightning to its Nemotron 3 family of open models, and it also released NeMo Switchyard, an open-source library for model routing. Lightning is a 30-billion-parameter mixture-of-experts model built with help from the Nemotron coalition. Nvidia says its reasoning is close to Nemotron 3 Super, even if both of its new models still trail Google’s Gemma 4 31B on Artificial Analysis’ Intelligence Index.

The company is leaning hard into a different pitch: speed, customization, and systems made of multiple models working together. In Nvidia’s telling, a larger frontier model can plan the job while a smaller, tuned model handles execution. Joey Conway, Nvidia’s senior director, said the company sees these systems as the future of AI.

Lightning is also meant to be fast in the literal sense. Nvidia says it can produce output up to 4x faster, and Kari Briski said the bigger advantage is that developers can modify and optimize it for specific workflows. She said general agentic benchmarks are only the starting point, because production work depends on task accuracy, and post-training can make a major difference.

Nvidia says early access customers used Lightning for specialized workflows and improved accuracy after post-training. Working with partners including CrowdStrike and CodeRabbit, it found that a fine-tuned open model like Lightning could match, and sometimes beat, larger proprietary models on the specific task it had been trained for. That matters because enterprise buyers are paying close attention to frontier-model costs, even if the extra tuning work is not trivial.

Switchyard is Nvidia’s answer to the routing layer underneath all of this. The library is written in Rust, but developers interact with it through APIs, so it is meant to fit into existing stacks. Users define a model pool and routing policies, then tune for quality, latency, and cost. Nvidia says systems using Switchyard with a mix of open models and Anthropic’s Opus 4.8 cut task-completion costs to about a third of running Opus alone in internal tests, while partner numbers from LangChain and Ramp showed similar savings. Lightning is available on Hugging Face, ModelScope, OpenRouter, and build.nvidia.com as an Nvidia NIM microservice, while Switchyard is on GitHub.

My take — AI-written commentary, not fact-checked reporting

This is Nvidia making a very sensible bet: the next AI fight is not one giant model, but who controls the router between many models. That’s less glamorous than a benchmark victory and a lot more useful. Open models plus routing is also the kind of boring enterprise story that actually sticks, which is usually where the money hides.

Read more about this at: The New Stack

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.