TLDRocket
Sign in

NVIDIA Nemotron 3.5 Lightning now available in Amazon SageMaker JumpStart

Amazon Web Services Venu Kanamatareddy

NVIDIA’s Nemotron 3.5 Lightning is now in SageMaker JumpStart. AWS says it’s built for always-on agents that need speed, not giant infrastructure.

Based on reporting by Amazon Web Services, Venu Kanamatareddy — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

AWS has added NVIDIA Nemotron 3.5 Lightning to SageMaker JumpStart, so teams can deploy the model without setting up the serving stack themselves. That’s the main selling point here: less plumbing, faster access, and a model aimed squarely at high-volume agent workloads.

Nemotron 3.5 Lightning is an open model, distilled from NVIDIA’s Nemotron 3 Ultra and built with the Nemotron Coalition. It uses a hybrid Mixture-of-Experts design, has 30B total parameters with 3B active on each forward pass, and supports up to 1M tokens of context. NVIDIA says that combination makes it a fit for repetitive agent steps that don’t need frontier-scale compute.

The performance claims are sharp. NVIDIA says Lightning can deliver up to 4x higher throughput and up to 30% faster task completion on high-volume agentic workloads. The article also says the published evaluations were run under a consistent harness, and that NVFP4 stays close to BF16 on several benchmarks, including MMLU Pro, GPQA Diamond, SWE-bench Verified, PinchBench, IFBench, and AA-LCR.

This is really a model-routing story as much as a model-launch story. AWS points to a system where a larger model can handle the hard planning work, while Lightning takes the busywork: alert classification, form extraction, policy checks, log queries, summaries, the kind of calls that pile up fast. If NVIDIA’s NeMo Switchyard is in the stack, it can route those steps across a model pool and pick Lightning when speed matters more than brute force.

For deployment, AWS says users can launch it from SageMaker Studio, from the Hugging Face model page, or with the SageMaker Python SDK. There are two variants in JumpStart, NVFP4 and BF16, and the launch note also warns that the endpoint costs money while it’s running. Delete it when you’re done. Cloud bills have a way of developing an agentic personality of their own.

My take — AI-written commentary, not fact-checked reporting

This is the sensible use of open models: let them do the repetitive work and stop paying frontier-model rent for every tiny step. The industry keeps pretending one giant model should do everything, then acts surprised when latency and cost show up like unwanted guests. Routing is the real product here, and the sooner teams admit that, the better.

Read more about this at: Amazon Web Services

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.