TLDRocket
Sign in

Nvidia just showed that the harness, not the AI model, is now the real hero

TechCrunch Julie Bort Covered by 4 sources

Nvidia says the agent’s harness matters more than the model for long tasks. With the right scaffolding, Claude Opus 5 hit 100% on a tough benchmark; without it, 30%.

Based on reporting by TechCrunch, Julie Bort — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Nvidia’s latest research pushes a blunt idea: when AI has to do a long job, the wrapper around the model can matter more than the model itself. On Friday, the company said its researchers got Claude Opus 5 to score 100% on ARC-AGI-3 by using a custom harness built to manage memory well and adding a supervisor component. Run the same model without that setup, and the score fell to 30%, which was still the best result among the models Nvidia tested.

That gap is the point. Long-horizon work is not a single prompt and a single answer; it’s the messy business of making many decisions in sequence, sometimes over days. Nvidia’s Adel El Hallack described the harness as the scaffolding around the model: the runtime, the tools, the skills, the memory, the context, the feedback. The model is part of the brain, but not the whole agent.

The benchmark Nvidia picked makes the result even more pointed. ARC-AGI-3 is a set of 2D games with no instructions, so the model has to figure out how to play and win. A perfect score means beating the games as well as humans. OpenAI has already struggled here, saying its own models scored below 10% and then showing in separate research that just changing two harness settings tripled those scores. Still, Nvidia says the supervisor layer was the bigger breakthrough, because it nudges the system when it wanders into dead ends.

That supervisor is not a brand-new idea, but Nvidia argues most users are still running a much thinner setup than they think. Tools such as Claude Code, Codex, and Hermes usually rely on one layer of harnessing, while Nvidia’s researchers built a more elaborate system called Agentic Variation Operators, or AVO. Nvidia says AVO is not a product, and that much of the company’s Nemo-branded harness tech is openly available.

There’s also a cost angle here. Databricks said in July that the harness can dramatically change what the same model costs to run, and that the wrong one can double the bill. Nvidia’s bigger message is pretty simple: open models are only part of the story. Open harnesses, open runtime, open infrastructure — that is where control, and a lot of the performance, actually lives.

My take — AI-written commentary, not fact-checked reporting

This is the part of agent hype that keeps getting glossed over: the model is the celebrity, but the harness does the actual work. OpenAI can keep selling the magic trick; the grown-up engineering is in the scaffolding, and that’s where the control freaks and the sober people should be paying attention. Closed stacks look clever until the bill and the failure modes show up at the same time.

Read more about this at: TechCrunch

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.