AivexaNewsSearch
AI news for builders and product teamsChecked every hour
NVIDIA Developer BlogFirst partyDeveloper tools

How Full-Stack NIM Optimizations Deliver 2.5x More Users on Nemotron 3 Ultra

Collected Sep 30, 2026

NVIDIA published details of full-stack serving optimizations it says let the Nemotron 3 Ultra NIM serve up to 2.5x more users on a 4xB200 system compared with a baseline serving stack without NIM optimizations.

The company reports the NIM 2.0.12 optimized serving stack achieves 1,997 tokens per second at a target of 50 tokens per second per user, a 20 ms inter-token latency. The benchmark definition used is 4xB200 hardware with an agentic workload of 64K/400/76% KV reuse.

According to NVIDIA, the measured gains come from interacting configuration bundles rather than independently additive switches. The listed optimization layers are autotuned mixture-of-experts and Mamba kernels mapping the hybrid architecture to Blackwell GPUs; tensor parallelism across four GPUs with expert-aware execution; prefix caching, partial-prefix matching and Mamba state-cache tuning for context reuse; scheduler, batching and GPU-memory allocation tuning; and MTP speculative decoding, whose incremental benefit depends on acceptance rate and available memory headroom.

NVIDIA describes NIM as packaging model- and GPU-aware serving choices into a deployable microservice with validated configurations, standard APIs and an enterprise container lifecycle through NVIDIA AI Enterprise. It says NIM Certified adds regular inference-stack updates, CVE handling, broader hardware validation and commercial support.

The published curves are described as a starting point rather than a promise of identical results. NVIDIA recommends benchmarking with its AIPerf tool, replaying representative traffic such as a Mooncake-format JSONL trace, pinning the image tag or digest for each run, and selecting the Pareto point meeting a latency SLO. For agentic workloads on a four-GPU B200 system, the company points to the profile vllm-nvidia-b200-nvfp4-tp4-pp1-throughput-90.0 with speculative decoding enabled. It says more performance-optimized NIM configurations are planned across a broader range of models.

The NIM is available for download on NGC, with a hosted API also offered.

Read at NVIDIA Developer Blog

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

Deploying a large language model is only the first step toward production-ready serving. Production teams also need to serve as many concurrent users as...