Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses
Microsoft Zhiyuan He, Yuqing Yang
Microsoft Research Asia built Agent Lightning v1.0, a tiny RL framework for training real agents. It works with the same harness used in deployment, so training looks a lot less fake.
Based on reporting by Microsoft, Zhiyuan He, Yuqing Yang — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Microsoft Research Asia has shipped Agent Lightning v1.0, a reinforcement learning framework built around a simple idea: stop rebuilding the agent just to train it. Instead, the same harness used in deployment stays in place, and the training system hooks into it through an LLM proxy. That means the agent that learns is much closer to the agent that later ships.
That shift matters because a lot of agentic RL has been forced into a brittle compromise. Traditional setups assume the training framework owns the whole loop, but real coding agents bring their own tool calls, context handling, execution logic, and dependencies. Recreating all of that inside a trainer is expensive, and it can change behavior in ways nobody wanted. Agent Lightning’s answer is to leave the harness alone and let the framework observe the model calls instead.
The project is also unusually small. Microsoft says the whole framework comes in at roughly 3,500 lines of code, split across an API gateway, a rollout controller, and a customized trainer built on verl. The gateway acts as an OpenAI-compatible proxy and records prompts, responses, and log probabilities. The rollout controller runs agents either locally or as standard Kubernetes jobs. The trainer then gathers the samples and assembles them for learning.
There’s a nice practical edge to the design. Agent Lightning v1.0 does not need paid commercial sandbox services to spin up agents at scale; it can use self-managed clusters, cloud Kubernetes, or local infrastructure. The researchers also describe a collocated async setup that shares GPUs between rollouts and updates, and say it delivered about a 2x end-to-end speedup over synchronous RL while using fewer GPUs than conventional asynchronous RL.
The clearest proof is a coding-agent run built on SWE-smith, mini-SWE-agent, and Qwen3.5-9B. Using about 6,000 training samples, the pipeline lifted Qwen3.5-9B from 41.8% to 56.4% Pass@1 on SWE-bench Verified, a gain of 14.6 percentage points. That’s the kind of result that makes all the plumbing arguments feel less academic and more like the actual product.
My take — AI-written commentary, not fact-checked reporting
This is the right direction: train the real harness, not a cosplay version of it. Too many AI teams still treat deployment as an afterthought and then act surprised when the model behaves differently in the wild. Open-source, reproducible, and annoyingly practical is a much better look than another giant framework pretending to be inevitable.
Read more about this at: Microsoft