AivexaNewsSearch
AI news for builders and product teamsChecked every hour
NVIDIA Developer BlogFirst partyDeveloper tools

Manage Kubernetes Node Fleets with NodeWright

Collected Sep 30, 2026

NVIDIA officially introduced NodeWright, an open-source, Kubernetes-native package manager for configuring and updating host operating systems across Kubernetes node fleets. The project has run in production at NVIDIA as Skyhook, according to the announcement, and this post introduces the NodeWright name.

The operator watches NodeWright custom resources and orchestrates a sequence on each targeted node: cordon, wait for critical workloads to finish, drain remaining pods, apply and configure packages, interrupt if needed, then uncordon. It respects PodDisruptionBudgets by default, with configurable drain behavior, and lets operators declare pods that must never be interrupted by label.

Packages are container images carrying scripts, configurations, and binaries, plus verification scripts that surface failures and stop rollouts. They can set sysctl and GRUB parameters, configure crash dump collection, create logical volumes, install security agents, and remediate CVEs. NodeWright tracks package state and semantic versions to distinguish fresh installs, upgrades, and downgrades, and packages can declare dependencies to determine execution order.

The DeploymentPolicy resource provides progressive rollouts across compartments, named groups of nodes selected by labels with their own disruption budget and strategy. Strategies are fixed batch size, linear increase by a fixed delta, or exponential growth by a factor. Each includes a batch threshold for minimum success percentage before advancing and an optional failure threshold that stops a compartment after too many consecutive failed batches.

NVIDIA publishes tuning and node setup packages, including an intent-based tuning model where operators declare their accelerator and workload and the package assembles kernel parameters, power management, and system settings. NodeWright integrates with NVIDIA AI Cluster Runtime to apply host-level parts of version-locked recipes, and works alongside NVCRE for pre-workload validation and NVSentinel for runtime fault monitoring and cordon, drain, and remediation. NVIDIA states NodeWright does not replace the NVIDIA GPU Operator or NVIDIA Network Operator and manages the host OS layer underneath them.

NodeWright installs with Helm as an OCI artifact and is licensed under Apache 2.0 as part of DSX OS.

Read at NVIDIA Developer Blog

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

Kubernetes manages what runs on your nodes. Managing the nodes themselves is the challenge: kernel settings, system packages, storage layouts, security agents,...