AivexaNewsSearch
AI news for builders and product teamsChecked every hour
NVIDIA Developer BlogFirst partyDeveloper tools

How SWE-Serve Exposes the Gap Between Local Tests and Live Serving

Collected Sep 30, 2026

NVIDIA introduced SWE-Serve, a benchmark for evaluating AI coding agents on changes to inference-serving software, developed with input from the SGLang team. It contains 53 executable tasks derived from 83 merged SGLang pull requests, spanning six inference-engineering families: model enablement, decoding, caching, scheduling, serving APIs, and runtime performance.

Twelve tasks run on CPU and 41 use a single NVIDIA H100. The first release does not evaluate other inference engines, multi-GPU execution, or multi-node serving. Thirty-seven tasks come from a single upstream pull request, while 16 combine two to six related changes. The median reference solution modifies 553 lines across seven files, and a typical verifier has seven tests for new behavior and 10 regression tests. Nineteen tasks start a real server and three enforce a calibrated performance gate on an H100.

Across the 19 live-serving tasks, the same 627 patches pass 45.9% of the time under the complete verifier, rising to 69.4% when live-serving tests are excluded — 147 patches changed from fail to pass. Those tasks contain 276 live-serving tests, 242 sourced or adapted from SGLang. In the Gemma 4 MoE task, 16 of 33 patches passed every other check but failed at least one live-serving test.

SWE-Serve also reports that 26 tasks confined to one runtime domain have a 69.0% pass rate, versus 47.7% for 27 tasks spanning multiple domains, a 21.3 percentage-point difference that appeared in every model setting. Eleven models and 31 model-effort configurations were evaluated with mini-swe-agent under closed-book conditions, with mean pass@1 from 34.6% to 75.5%. NVIDIA says a SWE-Serve pass means only that a patch satisfies the benchmark verifier, not that it is deployable or endorsed by SGLang maintainers.

Read at NVIDIA Developer Blog

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

An AI coding agent’s patch can pass tests yet fail when the server loads a real model and handles requests. Evaluating changes to inference-serving software...