TLDRocket
Sign in

The Sequence Chat - Issue 912: NVIDIA’s Chris Alexiuk Talks About Nemotron, GPUs and Agentic AI

Substack Jesus Rodriguez Covered by 2 sources

NVIDIA’s Chris Alexiuk says Nemotron is shifting from chat models to agent work. The new Lightning model is meant to do the boring grunt work, fast and cheap.

Based on reporting by Substack, Jesus Rodriguez — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

NVIDIA’s Nemotron line is getting more pointed. In this interview, Chris Alexiuk frames it less as a chatbot family and more as an attempt to build the parts of AI that developers actually have to ship: models, data, recipes, deployment patterns, and the plumbing around all of it.

That’s the thread running through Nemotron’s odd-looking evolution. The earlier 340B model was pitched as a synthetic data factory. Then came Llama Nemotron and the Nemotron-H hybrids. Now Nemotron 3 is being positioned as a frontier open family, but one that is clearly designed with hardware in mind. The Nano, Super, and Ultra tiers line up neatly with different deployment scales, and Alexiuk doesn’t pretend that’s an accident. Right-sizing models to real hardware, he says, is the practical move.

He makes the same argument for open weights and open data. NVIDIA, in his telling, is not trying to win the model race against closed labs. It wants developers, researchers, and companies to build on top of AI, and it thinks scrutiny from more people is part of the safety story. That’s why Nemotron ships not just weights, but also data, recipes, tech reports, cookbooks, and curation methods. Alexiuk says the company has released tens of trillions of tokens of data alongside the models, because without the data, you can’t really audit what the model learned.

The technical choices are just as pragmatic. Nemotron 3 blends Mamba-2, sparse MoE, and a limited amount of attention, with Alexiuk saying attention is still needed, especially for long-context recall. LatentMoE is there to improve latency and throughput while making room for higher top-k, and he says routing specialization is real even if the human-friendly story around it is simplified. The larger point is that NVIDIA keeps testing what survives at scale, even when some directions don’t make it to prime time.

The newest piece of the puzzle is Nemotron 3.5 Lightning, which Alexiuk calls the “muscle” for AI agents. It’s a 30B MoE model with 3B active parameters, built for tool calls, subagent delegation, and the other bits of agent work that shouldn’t burn a frontier model. NVIDIA also trained it for the agent harnesses people already use and optimized it for deployment from data center clusters down to a local DGX Spark using NVFP4 quantization. That’s the real story here: less demo magic, more infrastructure.

My take — AI-written commentary, not fact-checked reporting

This is the right kind of open model strategy: ship the weights, ship the data, ship the recipes, and let the adults inspect the machinery. The industry loves pretending safety comes from secrecy, which is a nice trick until the first real bug report arrives. NVIDIA is playing the long game here, and annoyingly, it looks like the sensible one.

Read more about this at: Substack

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.