AivexaNewsSearch
AI news for builders and product teamsChecked every hour
NVIDIA Developer BlogFirst partyDeveloper tools

Validate GPU Cluster Readiness Before AI Workloads Land

Collected Sep 30, 2026

NVIDIA introduced the NVIDIA Cluster Readiness Engine (NVCRE), an open source Kubernetes controller that validates GPU cluster readiness by running real distributed workloads across topology-aware node groups before production workloads land.

NVCRE addresses the problem that a GPU cluster can pass every health check and still fail to run an AI workload. A 512-GPU training job can underperform or fail due to a slow GPU, a link that degrades under load, or a configuration routing traffic over a slower path. Operators may not discover the problem until hours into a run or until a customer files a ticket.

The controller uses a layered API of Certification, Workflow, and Job custom resources. Each failure is attributed to a specific node and category, such as NCCL communication or NeMo pretraining. A built-in catalog covers five NCCL communication variants, the NVIDIA Data Center GPU Manager level-4 diagnostic suite, and NVIDIA NeMo pretraining with NVIDIA Nemotron 5 models at 8B and 56B parameters.

Adaptive fault isolation automatically splits failing groups and reruns tests until it identifies a small set of suspect nodes instead of implicating the entire group. Pass and fail criteria use Common Expression Language and are evaluated against measured metrics, with no thresholds shipped by default.

A separate WorkloadRun API handles repetitive multi-node GPU workload setup on Kubernetes, including platform detection, framework-specific runtime configuration, and optional gang scheduling via a gang-aware scheduler such as KAI Scheduler.

NVCRE integrates with NVIDIA AI Cluster Runtime for validated configuration and NVSentinel for continuous telemetry-driven health monitoring, forming the NVIDIA DSX OS operating layer. NVCRE records failed nodes and reasons but does not cordon, taint, or patch node conditions. The NVSentinel NVCRE Certification Monitor can translate failed certification results into health events, and configured NVSentinel policies can quarantine and drain nodes or trigger remediation.

NVCRE requires Kubernetes 1.29 or later, kubectl, Helm 3.x, and NVIDIA GPU Operator. It is licensed under Apache 2.0 and developed in the open. Installation uses the nvcrectl setup init command.

Read at NVIDIA Developer Blog

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

A GPU cluster can pass every health check and still fail to run an AI workload. Even when every GPU, network link, and pod reports healthy, a 512-GPU training...