How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering?
Apple Machine Learning Research
The paper tests autonomous MLE coding agents and finds that elaborate multi-agent and retrieval harnesses don’t improve results over a minimal-harness coding agent when using an equal time budget and the same frontier LLM backbone. It specifically reports that open-source state-of-the-art harnesses offer no advantages over a single-session minimal-harness baseline. As a result, the authors argue that the extra machinery layers are mostly redundant for current MLE benchmarks and that spending effort on hand-crafted harnesses around strong models gives poor returns.
Why it matters
Recent autonomous machine learning engineering (MLE) agents have made significant progress on public leaderboards. Often motivated by progress stagnation over long-horizon cycles and limited Large Language Model (LLM) primitives, modern MLE agents are deployed on top of increasingly elaborate machinery: multi-agent orchestrators, dedicated retrieval subagents, and more. While such harnesses expand, the use of more primitive but improved coding agents—where LLMs have direct access to the execution environment through read, write, and bash primitives—has received little attention in the field…
Related stories
Can LLMs Engineer Their Own Agent Harness? ByteDance Seed’s HarnessDev Says Only 34 of 64 Changes Generalize
MarkTechPost · 3 weeks ago ·
33
A primer on self-improving agent harnesses
Substack · 2 months ago ·
7
AI Harnesses Are the Software Shells That Turn AI Models Into Capable Agents
Trending Topics · 1 month ago ·
9