AivexaNewsSearch
AI news for builders and product teamsChecked every hour
NVIDIA Developer BlogFirst partyDeveloper tools

Benchmarking LLM Inference at Scale with AIPerf

Collected Sep 30, 2026

NVIDIA has introduced AIPerf, described as the designated successor to GenAI-Perf and a ground-up rewrite, for benchmarking LLM inference at scale. According to NVIDIA, AIPerf does not run on top of Perf Analyzer the way GenAI-Perf did, a change the company says explains its scaling behavior.

AIPerf uses a multiprocess architecture: worker processes generate load, separate record-processor services handle results, and coordination runs over ZMQ. NVIDIA says this structure prevents the client from becoming a bottleneck during high-concurrency benchmarking, unlike single-process architectures it describes as GIL-bound.

The tool supports more than 15 endpoint types, including chat, responses, NIM rankings and image generation, along with public datasets such as ShareGPT and trace replay formats from Mooncake, Baseten and WEKA AgentX. Arrival patterns include constant, Poisson and gamma with tunable burstiness, gradual ramping for concurrency and request rate, and synthetic distributions including vLLM/SGLang range-ratio for variable input and output sequence lengths.

AIPerf reports TTFT, ITL, request latency and output token throughput with percentile breakdowns (p25, p50, p75, p90, p95, p99) plus minimums, maximums, averages and standard deviations. When DCGM or pynvml is available, it also captures GPU power draw, utilization and memory consumption in the same run output.

NVIDIA's walkthrough used Qwen3-0.6B served through vLLM. Flags such as --synthetic-input-tokens-stddev 0 pinned workloads to exactly 128 input and 128 output tokens, while --extra-inputs min_tokens:128 and ignore_eos:true forced the model to emit 128 tokens. A second run used --arrival-pattern poisson with --request-rate 10, --synthetic-input-tokens-stddev 128, --output-tokens-stddev 32 and --random-seed 42; input sequence lengths ranged from 154 to 818 tokens around a 512-token mean. NVIDIA notes --streaming is required to measure TTFT and ITL. AIPerf also handles multi-node Kubernetes deployments, KV cache reuse warm-up, trace replay, prefix synthesis, custom datasets and concurrency sweeps, per NVIDIA.

Read at NVIDIA Developer Blog

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

You’re deploying a model on a system. It starts up, prompts are getting responses. Now the hard question: Is this fast? Your instincts might lead you to send...