TensorRT Edge-LLM Completes the MLPerf Edge Agentic Benchmark 6.4x Faster on Jetson AGX Thor
NVIDIA reported that TensorRT Edge-LLM ran Qwen3.6-27B on a single NVIDIA Jetson AGX Thor Developer Kit in the MLPerf Inference v6.1 Edge Agentic benchmark, achieving 52.33 tokens per second and completing all 1,007 turns of the performance workload in 24 minutes and 36 seconds. According to NVIDIA, that is 6.4x faster than the llama.cpp reference submission of 2 hours and 37 minutes, which used Qwen3.6-27B with Q4_K_M quantization. The submission ran in SingleStream mode on one Jetson AGX Thor Developer Kit with 128 GB of unified memory at the MAXN power mode.
The benchmark measures an OpenAI-compatible model endpoint in performance and accuracy phases. The performance phase replays recorded software-engineering agent trajectories across 20 conversations and 1,007 generated turns, with input length growing to approximately 23.5K tokens; Intersection of Union based inline accuracy is also measured. The accuracy phase uses Berkeley Function Calling Leaderboard v4 prompts with single-turn only and reasoning off.
NVIDIA attributed the result to NVFP4 quantization for weights and activations, including the language-model head, plus FP8 for the KV cache. It said KV cache and recurrent-state reuse served approximately 96% of prompt tokens from hot cache, with only about 0.5M of the total 13.6M prompt tokens prefilled across the turns. The runtime identifies reusable prompt prefixes and restores cached attention KV pages, recurrent state and partial KV-page state for the hybrid Qwen3.6 architecture.
Tree-based multi-token prediction was also used, with the MLPerf server configuration using 8 draft steps, top-2 candidates at each drafting depth, and a 16-node verification tree. NVIDIA said this delivered an additional approximately 40% decoding performance gain over linear MTP with 3 draft steps for this workload. The implementation is available on the TensorRT Edge-LLM release/0.9.1-mlpinf branch, and NVIDIA said developers can start from a published, calibrated Qwen3.6-27B NVFP4 checkpoint.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
AI agents are moving from cloud data centers to vehicles, robots, and other edge devices. Unlike a chatbot that answers a single prompt, an agent works through...