AivexaNewsSearch
AI news for builders and product teamsChecked every hour

Heaps do lie: debugging a memory leak in vLLM.

Collected Oct 5, 2026

Mistral AI published an engineering account of investigating a suspected memory leak in vLLM, written by Mathis Felardos. The issue appeared during pre-production testing of disaggregated serving with one of its frontier models: memory usage climbed steadily under specific conditions, with vLLM, the model Mistral Medium 3.1, and graph compilation enabled. There were no crashes or errors, only a slow linear increase in system memory of 400 MB per minute on production-like traffic, which after a few hours would lead to an out-of-memory state.

The team could not reproduce the problem on a smaller model or with fewer production optimizations. It was present only on a Prefill/Decode disaggregated setup with NIXL. In that setup, a router sends a prefill request to a prefill vLLM instance to compute the KVCache, then transfers KVCache metadata with a decode request to a decode vLLM instance; transfer is initiated through NIXL and token generation happens on the decode instance. The leak was observed only on the decode side. NIXL relies on UCX (Unified Communication X), including over Infiniband.

Python memory profilers Memray and Guppy 3 showed no leak. GDB made the process crash, and Valgrind was impractically slow for the vLLM setup. The team opened a GitHub issue with the vLLM project, which helped confirm others were seeing the issue.

Heaptrack traced malloc and free but showed only a lazy NIXL initialization increase; peak RSS still differed between two snapshots taken around a benchmark. The team then used pmap to inspect /proc memory maps, finding that some anonymous memory mappings grew over time with changing start addresses. BPFtrace tracing showed the growing addresses came from mmap calls, not mremap, originating from glibc's raw syscall wrapper (syscall+29). LD_PRELOAD hooks did not intercept them. Switching from Ubuntu 22.04 LTS to Ubuntu 24.04 LTS, which enables frame pointers for native libraries, was not enough. The team identified UCX and PyTorch as potential culprits and turned to GDB automation, setting a conditional breakpoint on the syscall address when the call number matched SYS_mmap.

Read at Mistral AI

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt