Build Local AI Apps with C++ and NVIDIA TensorRT RTX Samples

NVIDIA has published Do Inference Now (DIN) Deploy, an open-source collection of C++ samples that combines ONNX Runtime with the NVIDIA TensorRT RTX execution provider to accelerate local AI inference on Windows and Linux. The same ONNX Runtime API can also be accessed through WinML 2.0.
Each sample separates model conversion from deployment: a Python exporter downloads a model checkpoint from Hugging Face and converts it into an ONNX artifact, while the application side is a native C++ CLI built on ONNX Runtime session and tensor APIs. Most sample code uses those ORT APIs; vendor-specific code, including CUDA APIs and kernels, appears only in optional accelerated paths.
The repository supports automatic speech recognition with OpenAI Whisper for offline transcription and with NVIDIA Parakeet TDT and NVIDIA Nemotron ASR Streaming for streaming pipelines. It also supports interactive masking for images and video with Meta SAM 2.1, and prompt-driven image generation with FLUX.2-klein-4B.
The FLUX.2 sample uses ONNX Runtime's graphics interop capability, introduced in version 1.25, with Vulkan and DirectX for sampling. It also shows how post-training quantization with NVIDIA Model Optimizer produces a quantized ONNX model that is a drop-in replacement requiring no application-code changes. CMake presets cover Windows and Linux, including Arm64 variants; DirectX is available only on Windows.
Performance measurements on DGX Spark show GPU acceleration ranging from 39x real-time for Nemotron ASR streaming to 206x for Parakeet TDT, while SAM 2.1 achieves 38.3 FPS on GPU versus 0.5 FPS on CPU. CMake downloads ONNX Runtime and TensorRT RTX by default.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
Adding AI models to local applications requires a portable model format, a reliable runtime, and acceleration that works across target systems. Do Inference Now...