Microsoft Releases Agent Lightning to Train AI Agents in Their Existing Setup
The lightweight framework keeps the deployment harness in the training loop. Microsoft reports better coding scores and faster training, but preserving real agent behavior creates its own technical challenges.
Agent Lightning v1.0, open-sourced by Microsoft Research Asia on October 7, 2026, trains agents through their existing harnesses: an OpenAI-compatible proxy captures model traffic while agent execution stays outside the trainer. In a specific coding pipeline using Qwen3.5-9B, Microsoft reports Pass@1 on SWE-bench Verified rose from 41.8% to 56.4% after about 6,000 training samples; the result does not establish gains across all supported harnesses. The framework also supports Kubernetes or local runs and sharing GPUs between agent runs and model updates.
01
Developers can usually connect an existing harness by changing its model endpoint, rather than rebuilding its tool and execution setup for training.
02
The roughly 3,500-line framework comprises an API gateway, rollout controller, and customized trainer; it can run without paid commercial sandbox services.
03
Context summaries and subagents can split one run into multiple training samples, potentially giving fragmented runs extra weight if samples are simply averaged.
Developers can now train AI agents while keeping the software that manages their tools, context, and execution intact. Microsoft Research Asia announced and open-sourced Agent Lightning v1.0 on October 7, 2026. The framework routes model calls through a proxy, allowing the agent’s deployment setup to participate directly in reinforcement learning rather than be rebuilt for training.
An agent harness is the surrounding software that controls tools, context, and the steps an agent takes. Reinforcement learning teaches systems through trial and error, using rewards and penalties. Microsoft’s premise is that rebuilding a harness inside a training framework is expensive—and can leave developers training something that behaves differently from the deployed agent.
Agent Lightning instead inserts an OpenAI-compatible model proxy between the harness and the model. Microsoft says changing the model endpoint is usually enough to connect an existing harness. The framework contains roughly 3,500 lines of code across three components: an API gateway, a rollout controller, and a customized trainer.
The gateway records prompts, responses, and the model probabilities needed for training. The controller launches agent runs, while the trainer collects their resulting samples. Agent execution stays separate from the trainer, instead of handing the training framework control of the agent’s interaction loop.
Simply averaging over those samples can give extra weight to runs that produce more of them, even when that count reflects harness behavior. Other challenges include text-to-token conversion and scheduling variable workloads on fixed GPU resources.
The coding pipeline combines SWE-smith, mini-SWE-agent, and Qwen3.5-9B, with data cleaning, environment construction, and safeguards against gaming the reward. Microsoft reports Qwen3.5-9B improving from 41.8% to 56.4% Pass@1 on SWE-bench Verified—a 14.6-percentage-point gain—using about 6,000 training samples. That result applies to the specified pipeline, not every supported harness.
Agents run as standard Kubernetes jobs on self-managed or cloud clusters, or on local systems, without requiring paid commercial sandbox services. Agent Lightning also lets agent runs and model updates share GPUs. When enough runs finish, the gateway pauses new requests, waits for active requests to complete, and resumes agent work after the update. Microsoft reports about twice the end-to-end speed of synchronous training in its experiments.
Diagram of Agent Lightning’s system architecture, with harness agents connected to the API gateway, rollout controller, and customized trainer.Source: microsoft.com.
Sources
arxiv.orgAgent Lightning v1.0: Towards Harnessed Agentic RL
Reader comments
Newest comments first. Replies stay oldest first.