Enabling Private High-Performance Production AI Inference with NVIDIA Confidential Computing
NVIDIA published a developer blog post describing how TensorRT LLM applies Confidential Computing (CC)-aware adaptations for production AI inference on Blackwell GPUs.
NVIDIA said CC runs workloads using memory-encrypted confidential virtual machines, confidential GPUs, and encrypted NVLink. It said secure execution changes assumptions around memory movement, timing, scheduling, and multi-GPU communication, and that such changes introduce performance overhead if the runtime does not adapt.
In a controlled comparison using eight NVIDIA B200 GPUs with the DeepSeek-R1 model, NVIDIA reported that CC on retained 96.1% to 98.2% of CC off output-token throughput, while mean time per output token stayed within 1.2% to 4.3% of baseline across concurrency levels 1 through 16. The comparison held model, hardware, framework version, sequence lengths, parallelism, and concurrency constant.
NVIDIA described three adaptations. On host-to-device transfers in B200 CC, data passes through a software encrypted bounce buffer because the GPU cannot directly access protected CVM memory; TensorRT LLM selects pageable memory for affected paths instead of pinned memory. For device-to-host, it moves repeated token and sampling-data readback to an asynchronous worker to avoid blocking the main scheduler during decode.
The kernel autotuner normally uses CUDA events, but NVIDIA said CUDA-event timestamps produced an unstable timing signal in the tested CC configuration, which could cause a slower tactic to be selected. TensorRT LLM uses the GPU %globaltimer for tactic measurements under CC while retaining CUDA events outside CC.
NVIDIA also said NVLS multicast is unavailable in B200 CC configurations, so frameworks should detect NVLS availability and choose communication algorithms that minimize latency for the given message size and topology. It said NCCL_SYMMETRIC cannot provide its intended multicast benefit without NVLS but may still incur memory registration and cross-rank synchronization costs.
NVIDIA recommended teams attest the environment and benchmark CC-on and CC-off using the exact workload they intend to serve.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
As large language model (LLM) inference increasingly processes sensitive information and proprietary model context across personal, enterprise, and regulated...