AivexaNewsSearch
AI news for builders and product teamsChecked every hour

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Collected Oct 7, 2026

Microsoft Research Asia has introduced a training paradigm it calls Harnessed Agentic RL and released a fully rebuilt, open-source Agent Lightning v1.0 around it. In this setup, the agent harness used in deployment takes part directly in reinforcement learning rather than being reimplemented inside the training framework. The release is positioned around three properties: a small codebase, integration with real harnesses, and a complete, reproducible agent RL training pipeline.

The framework is about 3,500 lines of code in total. Microsoft describes that size as deliberate, so the codebase stays small and clear enough to understand, modify and extend. Existing harness code is left unchanged: agents reach the model through an LLM proxy, and pointing the endpoint that previously called the model API at Agent Lightning is usually enough to connect a harness to RL training. Agents run as standard Kubernetes jobs on self-managed clusters, cloud Kubernetes or local infrastructure, with no dependency on paid commercial sandbox services.

The headline result comes from an end-to-end coding agent pipeline. Researchers built it on SWE-smith, mini-SWE-agent and Qwen3.5-9B, covering data cleaning, environment construction, reward-hacking safeguards and RL training. The training set holds roughly 6,000 samples and needs no large-scale compute. RL training alone raised the model from 41.8% to 56.4% Pass@1 on SWE-bench Verified, an absolute gain of 14.6 percentage points.

The motivation is a mismatch between how RL systems were built and how agents now work. Traditional agentic RL assumes the training framework owns the interaction loop with the environment. In a ReAct-style loop, the model generates an action, the environment returns an observation, the observation is appended to the context, and the model generates the next action, so the whole rollout maps onto one continuous token trajectory. Early systems such as verl, AReaL and slime were built this way, which meant training an agent required rebuilding its loop inside the RL framework.

Real harnesses, in Microsoft's account, have outgrown that assumption. Coding agents such as mini-SWE-agent, OpenHands, OpenCode, Claude Code and Codex each bring their own context management, tool protocols, execution logic and dependencies. Rebuilding one for training is expensive, and the rebuilt agent may no longer behave the same way as the deployed one. Agent Lightning instead sits an LLM proxy between the agent and the model. The agent runs as before, and the training framework observes and records its model calls.

Moving the loop into the harness creates four specific technical problems that the release spells out. First, retokenization and sample merging: harnesses keep context as text, but RL training needs the token IDs sampled during the rollout, and passing text through the chat template and tokenizer again can shift token boundaries, so adjacent calls cannot always be merged into one sample. Second, advantage calculation: retokenization, subagents and context summarization can split one rollout into several samples, and computing baselines and advantages directly at the sample level causes rollouts that produce more samples to be counted repeatedly, altering the statistical relationships at the rollout level. Third, loss normalization: averaging loss by sample count gives more weight to rollouts that produce more samples, and since sample count is often just a product of harness behavior, normalization must avoid being distorted by it. Fourth, training backend scheduling: sample count and length are known only after the harness finishes, while GPU counts and data or tensor parallel configurations are usually fixed, so the backend has to map a variable workload onto fixed resources.

The system has three core components. The API Gateway stores rollouts, models and events and serves as an OpenAI-compatible LLM proxy; it links every model call from the harness to its rollout and records the prompts, responses and log probabilities training needs. The Rollout Controller starts and manages agent execution, either as local processes or as standard Kubernetes jobs, keeping agent execution separate from the trainer. The Customized Trainer, built on verl, creates rollouts, waits for them to finish, collects samples, and assembles final training samples through a sample adapter.

Scheduling is handled by what the release calls Collocated Async RL. Rollout times vary widely across agents; synchronous RL waits for the slowest agent in a batch and leaves GPUs idle, while fully asynchronous RL raises utilization but needs separate GPU pools for rollout and training. Collocated Async RL lets rollout and model updates share the same set of GPUs. Once enough rollouts are collected, the API Gateway pauses accepting new requests and waits for in-progress requests to finish; rollout resumes after the update completes. Microsoft reports the state transition is transparent to the external harness. In experiments, this delivered about a 2x end-to-end speedup over synchronous RL while using fewer GPUs than conventional asynchronous RL.

On infrastructure, the argument is cost and portability. Collecting enough rollouts means running many agents at once, consuming substantial CPU, memory and compute. Other Harnessed Agentic RL frameworks often host agents on commercial sandbox services such as Modal Sandbox or E2B, where cost climbs quickly with scale, according to the release. Agent Lightning v1.0 runs them as standard Kubernetes jobs instead, reusing existing clusters or local infrastructure so large rollouts cost less and the pipeline stays open source and reproducible.

The experiments also served as a check on the advantage and loss questions. Compared with sample-level handling, rollout-level advantage combined with rollout-level normalization achieved a higher validation reward and kept policy entropy more stable during training on the SWE-smith validation set.

For a developer reading this, the practical appeal is the deployment-training gap. When an agent is rebuilt inside an RL framework, the trained policy and the shipped agent are different programs, and any behavioral difference from context handling or tool protocols is unlearned. Harnessed Agentic RL narrows that gap by construction, since the harness that trains is the harness that ships. The trade-off is that the training system sees only request-response pairs rather than a token-level environment loop, which is exactly why the four problems above exist and why rollout-level statistics matter. A likely second consideration is operational: Kubernetes-based rollouts remove per-sandbox billing but put cluster provisioning, isolation and cleanup on the team running training.

Why it matters: teams building coding or general-purpose agents now have an open-source route to fine-tune the exact harness they deploy, rather than a rebuilt approximation, and the reported 14.6-point gain from about 6,000 samples suggests meaningful improvement is reachable without large-scale compute. The Kubernetes-native design means infrastructure already in place can be reused, though sandboxing and scheduling become the team's responsibility. Developers evaluating it should treat the SWE-bench Verified result as specific to the Qwen3.5-9B, SWE-smith and mini-SWE-agent pipeline described, and verify transfer to their own harness.

Read at Microsoft Research

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research .