Simplifying Model Serving Across Multiple GPUs with NVIDIA TensorRT Multi-Device Integration in NVIDIA Dynamo-Triton
NVIDIA Dynamo-Triton release 26.07 enables the multi-device capability of the TensorRT backend, which lets a single Triton KIND_MODEL instance own multiple GPUs, create per-rank TensorRT execution contexts, CUDA streams, and NCCL communicators, and launch the ranks together per request. Applications call one named model through a gRPC endpoint rather than coordinating GPU ranks themselves. TensorRT multi-device inference is fully supported starting with TensorRT 11.0.
The integration is demonstrated with NVIDIA Cosmos 3 Nano video generation. Diffusers orchestrates prompts, latents, classifier-free guidance, scheduling, VAE decode, and frame postprocessing, while Dynamo-Triton serves the 36-layer denoising transformer. TensorRT multi-device inference uses Ulysses context parallelism to distribute 44,160 video tokens across as many as eight GPUs. The distributed Ulysses graph is compiled into each TensorRT plan before deployment; Dynamo-Triton activates a context-parallel plan rather than converting a single-device engine.
Each of 35 denoising steps requires a negative or unconditional prediction and a prompt-conditioned prediction for CFG, so the Diffusers proxy makes two sequential Triton calls per step, for 70 transformer RPCs per generation.
All four variants ran on the same eight-GPU system, with 1280x720 output, 189 frames at 24 FPS, and 35 denoising steps; each result includes one warm-up and five measured generations. End-to-end latency dropped from 156.595 seconds on one GPU to 34.183 seconds on eight GPUs, with transformer RPC speedup reaching 6.09 times. Transformer RPCs accounted for 93.4% of generation time on one GPU and 70.2% at CP8.
Every variant used the same seed and profile. Validation sampled frames 0, 47, 94, 141, and 188. CP2, CP4, and CP8 passed configured thresholds of mean absolute error at or below 25 and PSNR at or above 18 dB, with CP2 and CP4 measuring MAE 12.759 and PSNR 21.111 dB, and CP8 measuring MAE 16.316 and PSNR 19.400 dB. Outputs are not claimed to be pixel-identical. The benchmark does not measure concurrent request throughput, cost per generated video, or total cost of ownership.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
The compute and memory demands of generative AI increasingly exceed what a single GPU can provide. NVIDIA TensorRT multi-device inference is a new capability...