NVIDIA Nemotron 3.5 Lightning and NeMo Switchyard Deliver Faster, Smarter, More Efficient Agentic AI
NVIDIA Kari Briski ● Covered by 4 sources
NVIDIA launched Nemotron 3.5 Lightning and an open routing tool for AI agents. The pitch: faster specialist models, less cost, and more control over where the work runs.
Based on reporting by NVIDIA, Kari Briski — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
NVIDIA is pushing harder into the agent era with two releases aimed at the plumbing, not the chatbot demo. Nemotron 3.5 Lightning is the newest member of its Nemotron 3 family, and NVIDIA says it is the highest-efficiency model in its class for long-running agentic workloads. The other piece, NeMo Switchyard, is an open source routing library that can steer requests to the best model inside an agent workflow.
The idea behind Lightning is simple: not every part of an AI system needs a giant reasoning model. NVIDIA frames modern agents as systems of models, where a larger frontier model like Nemotron 3 Ultra or GPT-5.6 can plan the work, while smaller specialists handle things like code review, tool use, security alerts or billing questions. Lightning is a 30-billion-parameter mixture-of-experts model built for those narrower jobs, and NVIDIA says it can produce output up to 4x faster, with agentic task completion 30% faster than other models in its class.
Because it is open and customizable, the model can be post-trained with NVIDIA NeMo on an organization’s own data, tools and workflows. NVIDIA says that should help with accuracy on specialized tasks, and it points to customers and partners already adapting it for cybersecurity, legal services, code review, software development, finance, healthcare and even physical and life sciences. The company also says Lightning can run locally or on premises, including on NVIDIA RTX PCs, DGX Spark, DGX Station and Jetson, or scale out across workstations, data centers and cloud environments.
NeMo Switchyard is meant to make the whole setup less clumsy. Instead of routing everything through one default model or manually stitching together model choices, it can direct each request automatically based on quality, latency and cost targets. NVIDIA’s internal benchmark says that approach keeps frontier-level accuracy while cutting task completion cost to nearly one-third of Opus 4.8 alone. Partners are already claiming real-world gains too, from Boomi’s 100% domain-routing accuracy to Ramp’s 58% cost cut and 33% runtime drop in Ramp SWE-Bench.
Both tools fit NVIDIA’s familiar playbook: give enterprises control, then wrap that control in enough infrastructure that it feels practical. Lightning is available through Hugging Face, ModelScope, OpenRouter and build.nvidia.com, plus NVIDIA NIM and partner platforms. Switchyard is on GitHub now, with partner platforms promised later.
My take — AI-written commentary, not fact-checked reporting
This is the part of the AI race that actually matters: who controls the model mix, not who has the loudest demo. Open models win when they let enterprises route work sensibly and keep data where it belongs; closed one-size-fits-all systems start looking expensive very quickly. The industry keeps selling intelligence, but the real product is control with fewer headaches.
Read more about this at: NVIDIA
Related stories
NVIDIA AI Releases Nemotron 3.5 Lightning: A 30B Open MoE with 3B Active Parameters, and NeMo Switchyard Model Router
MarkTechPost · 3 weeks ago ·
9
NVIDIA Nemotron Achieves Benchmark-Leading Performance With LangChain Deep Agents Harness
NVIDIA · 1 month ago ·
51
The Sequence Chat - Issue 912: NVIDIA’s Chris Alexiuk Talks About Nemotron, GPUs and Agentic AI
Substack · 3 weeks ago ·
5