Microsoft Research Asia Releases Agent Lightning v1.0: A 3,500-Line Framework for Harnessed Agentic RL
Key point
Microsoft Research Asia released Agent Lightning v1.0, a 3,500-line framework that raised Qwen3.5-9B's Pass@1 on SWE-bench Verified by 14.6 percentage points using only 6,000 training samples.
Details
Microsoft Research Asia has introduced Harnessed Agentic RL, a training paradigm where the deployment agent harness participates directly in reinforcement learning, eliminating the need to reimplement agents inside training frameworks. The researchers open-sourced Agent Lightning v1.0, a lightweight control plane built on this paradigm, designed to integrate with real-world agent systems without modifying their existing code.
Lightweight Architecture and Kubernetes Support
Agent Lightning v1.0 is implemented in approximately 3,500 lines of code, focusing on simplicity and reproducibility. The system comprises three core components: an API Gateway that acts as an OpenAI-compatible LLM proxy to record model calls, a Rollout Controller that manages agent execution, and a Customized Trainer built on verl.
The framework features native Kubernetes support, allowing agents to run as standard Kubernetes jobs on self-managed clusters, cloud infrastructure, or local machines. This approach avoids dependency on paid commercial sandbox services, reducing costs and improving resource efficiency for large-scale rollouts.
Performance and Training Efficiency
The team introduced Collocated Async RL, a method that allows rollout and model updates to share the same GPU resources. This technique achieved approximately a 2x end-to-end speedup over synchronous RL while using fewer GPUs than conventional asynchronous methods.
In a practical demonstration, researchers built an end-to-end coding agent pipeline using Qwen3.5-9B and mini-SWE-agent. Using only about 6,000 training samples from the SWE-smith dataset, the RL training raised the model's Pass@1 score on SWE-bench Verified from 41.8% to 56.4%, an absolute gain of 14.6 percentage points. The experiments also validated that rollout-level advantage calculation and loss normalization provide more stable training compared to sample-level handling.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.