AivexaNewsSearch
AI news for builders and product teamsChecked every hour

Control How Your GPU Shares Work with Green Contexts

Collected Oct 6, 2026

NVIDIA detailed green contexts, a CUDA feature that lets an application explicitly select a subset of GPU execution resources and direct work to those resources. Green contexts have been available in the Driver API since CUDA 12.4; starting with CUDA 13.1 they are also accessible through the Runtime API, enabling an application, within its process, to define where work runs and how execution resources are divided.

According to NVIDIA, controlling how GPU resources are shared between concurrent components remains difficult: components can interfere unpredictably, and existing tools offer limited partitioning ability. Traditional CUDA contexts were not designed for this usage model, NVIDIA states, describing them as heavyweight, incurring hardware context-switch overhead, and reflecting assumptions from an era when GPUs were smaller and applications typically ran as a single dominant workload.

One primary use is SM partitioning, NVIDIA says, where assigning a specific subset of SMs to a green context targets work submitted through it to those SMs, allowing multiple workloads to run concurrently without competing for the same compute units. Green contexts can also provision workqueue resources; NVIDIA says independent stream-ordered workloads may otherwise map to the same underlying workqueues, introducing unintended serialization even when execution resources are sufficient.

NVIDIA describes green contexts as lightweight to create and destroy, with creation or destruction not implicitly synchronizing unrelated GPU work. In the Runtime API, they are represented by the cudaExecutionContext_t type; cudaGreenCtxCreate() returns a handle that can be passed to APIs such as cudaExecutionCtxStreamCreate() instead of relying on implicit thread-local device or context state. Applications targeting the full device can continue using the traditional Runtime model, or use cudaDeviceGetExecutionCtx() for APIs taking an explicit context handle.

NVIDIA's example covers a latency-sensitive kernel sharing a device with a throughput-oriented bulk kernel, such as communication/GEMM overlap or latency-sensitive operators in NVIDIA Holoscan. It notes stream priority alone cannot preempt a block already executing on an SM. On an NVIDIA Blackwell GPU with 148 SMs, NVIDIA tested three modes with the same critical and bulk workload: green-context partition with a high-priority critical stream; default context with a high-priority critical stream but no partition; and default context with normal priority on both streams. NVIDIA reports stream priority alone was roughly 27x faster than equal-priority streams competing for the same SMs, and that an additional 20x was spent waiting for bulk blocks to drain. The tradeoff, NVIDIA states, is fewer SMs available for the bulk kernel.

NVIDIA advises building with CUDA 13.1 or newer, querying device resources, creating a green context for the desired partition, then creating streams for it. Green contexts are opt-in and additive, so existing applications can target the full device unchanged.

Read at NVIDIA Developer Blog

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

GPU applications increasingly consist of multiple independent components running at the same time within a single process: a latency-sensitive operator...