AivexaNewsSearch
AI news for builders and product teamsChecked every hour

How DOCA GPUNetIO Unifies GPU-Initiated Networking Across the NVIDIA Software Stack

Collected Oct 6, 2026

NVIDIA detailed how DOCA GPUNetIO serves as a unified GPU-initiated networking foundation across its software stack. The framework allows CUDA kernels to directly drive Ethernet, RDMA, Verbs, and DMA operations while keeping the CPU out of the application critical path, according to the company's developer blog.

DOCA GPUNetIO provides CPU functions on the control path to export transport objects created with DOCA Ethernet, DOCA Verbs, DOCA DMA, and DOCA CommChannel, and GPU CUDA functions on the data path so kernels can manipulate those exported objects.

NVIDIA ships GPUNetIO in two forms: a full DOCA SDK version that is a superset spanning Verbs, Ethernet, DMA, and Comm Channel integration, and a lighter-weight open-source project focused on RDMA Verbs. NVIDIA said the open-source implementation can detect the DOCA SDK at runtime and call selected closed-source functions through dlopen; without the SDK it runs standalone.

Multiple communication libraries now build on the shared GPUNetIO foundation rather than maintaining separate GDA-KI implementations, including NCCL, NVSHMEM, UCX/NIXL, and Holoscan Sensor Bridge. NVIDIA said NCCL GIN has integrated the open-source GPUNetIO Verbs path as a backend since version 2.27, and NVSHMEM 3.7 introduced a GPUNetIO-based transport that NVIDIA says reduces implementation complexity while preserving IBGDA performance. Benchmarks cited by NVIDIA show GDA-KI enables better CTA and QP scaling for small message sizes.

NVQLink uses the GPUNetIO-based GPU RoCE Transceiver operator to achieve approximately 2.6 microseconds minimum round-trip latency for quantum-classical workflows on IGX Thor with a Blackwell GPU and ConnectX-7, NVIDIA said. Other integrations include the Aerial 5G SDK, Holoscan Advanced Network Operator, and DeepEP/HybridEP.

The framework exposes high-level and low-level APIs for Ethernet and RDMA Verbs transports. High-level functions such as doca_gpu_dev_verbs_put_signal and doca_gpu_dev_eth_txq_send combine multiple operations and handle concurrent WQE submission and doorbell ringing, while low-level primitives let applications build custom composite operations and manage synchronization themselves. Doorbell options include regular MMIO mapping, BlueFlame for latency-sensitive applications, and a CPU-assisted mode for systems without a direct GPU-to-NIC connection such as DGX Spark.

NVIDIA pointed developers to the DOCA SDK GPUNetIO programming guide, the open-source repository, and the NCCL tests GIN device API for reference code.

Read at NVIDIA Developer Blog

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

GPU applications increasingly need networking and data movement to behave like first-class GPU-controlled operations rather than host-driven services. When the...