ThunderAgent: 2x Faster Agentic Inference for Synthetic Data Generation at Scale

Together AI, working with researchers at Georgia Institute of Technology, the University of Illinois Urbana-Champaign, and Carnegie Mellon University, introduced ThunderAgent, a scheduling system for high-throughput agentic inference. The work was accepted to ICML 2026 as a Spotlight paper.
ThunderAgent is a lightweight scheduling layer positioned between agentic clients and inference backends. It abstracts each agentic workflow as a schedulable program, tracking execution phase, KV cache footprint, and node placement. The system addresses KV cache thrashing, which occurs when request-level engines evict an agent's cache while it waits on a tool call and must recompute the conversation on resume. ThunderAgent monitors per-node memory pressure, selectively pauses low-priority workflows, and routes resumed workflows through a global waiting queue to the node with the most available capacity.
In the company's internal synthetic data generation pipeline, tested on a single 8xH100 node with HiCache offloading, ThunderAgent reached 803 token/s throughput with mean latency of 10.6s at batch size 192, compared with 390 token/s and 65s mean latency for SGLang's default scheduler. On multi-node clusters, ThunderAgent's throughput grew from 671 to 2,248 steps/min when scaling from 16 to 64 GPUs, with speedup over SGLang Gateway widening from 1.79x at 2 nodes to 2.39x at 8 nodes.
ThunderAgent interfaces with inference backends through OpenAI-compatible endpoints and requires only adding a program_id field client-side. Together AI says it is compatible with optimizations such as quantization and speculative decoding, and that it has been adopted by open-source frameworks including SkyRL and NVIDIA Dynamo. The project is open source, with a paper on arXiv.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
ThunderAgent is a program-aware scheduler for agentic inference. By treating each agent workflow as a schedulable program, it eliminates KV cache thrashing to deliver more than 2x single-node throughput and near-linear multi-node scaling.