TLDRocket
Sign in

NVIDIA Nemotron Achieves Benchmark-Leading Performance With LangChain Deep Agents Harness

NVIDIA Adel El Hallak

NVIDIA and LangChain tuned an open AI agent setup that matches top closed models on real tasks, but costs 10x less to run. They did it without retraining the model at all — just smarter engineering around it.

Based on reporting by NVIDIA, Adel El Hallak — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

NVIDIA is making a pointed argument with its latest Nemotron release: you don't always need a better model, you need a better harness. Working with LangChain, whose agent tooling gets over 200 million downloads a month, NVIDIA tuned the Deep Agents framework specifically around Nemotron 3 Ultra. The result, according to benchmarks run on LangChain's own Deep Agents suite, is an open model that beats every other open competitor on accuracy and matches the top closed models on business-task performance — while running at roughly a tenth of the inference cost.

What makes this notable isn't the raw score. It's how they got there. LangChain's engineers didn't retrain Nemotron 3 Ultra at all. Instead, they dug into execution traces from failed benchmark runs, found where the agent was tripping up, and fixed the environment around it — system prompts, tool descriptions, middleware. Harrison Chase, LangChain's cofounder and CEO, frames this as a philosophy shift: memory, tool use, evaluation and model behavior compound when a team tunes them together, rather than treating the model as the only lever worth pulling.

The cost angle is where this actually gets interesting for enterprises. At a tenth of the price of leading closed models, teams can afford to run evaluations constantly, iterate faster, and spin up specialized agents across far more of the business than a per-call budget would normally allow. NVIDIA is bundling this work into something called NemoClaw for LangChain Deep Agents — an open reference blueprint pairing the tuned harness with NVIDIA's OpenShell runtime for executing agent actions safely. Open model, open harness, open runtime: NVIDIA's pitch is that enterprises get to own the whole stack instead of renting it.

Companies like Abridge, Amdocs and Box are already embedding specialized agents built this way into their platforms, and EY is expanding its NVIDIA implementation work specifically around NemoClaw to help clients govern these systems in higher-stakes workflows. That last detail matters. As agents move from answering questions to taking actions inside core business systems — approving claims, moving money, editing records — who controls the stack, and how cheaply you can audit and retune it, stops being a nice-to-have and becomes the whole ballgame.

The tuned harness is live now through LangChain, and developers can access Nemotron 3 Ultra on Baseten, Crusoe Cloud, DeepInfra, Fireworks, Nebius and Together AI. NVIDIA clearly wants this treated less as a benchmark win and more as a template: proof that open stacks, engineered carefully, can close the gap with closed frontier models without anyone touching the weights.

My take — AI-written commentary, not fact-checked reporting

This is the most convincing argument I've seen yet for open models over closed ones in production: not

Read more about this at: NVIDIA

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.