AivexaNewsSearch
AI news for builders and product teamsChecked every hour
NVIDIA Developer BlogFirst partyDeveloper tools

Topology-Aware Workload Scheduling with NVIDIA Topograph

Collected Sep 30, 2026

NVIDIA introduced Topograph, an open source toolkit that identifies a cluster's network topology so workload managers can make topology-aware scheduling decisions. It discovers topology from cloud APIs or on-premises fabric systems, normalizes it into a canonical model, and publishes it in the format each workload manager expects: Kubernetes node labels, Slurm topology configuration, or Slinky ConfigMaps.

Topograph has two concepts: providers and engines. A provider discovers topology and normalizes it; an engine translates that model into Slurm configuration, Kubernetes labels, Slinky ConfigMaps, Node Feature Discovery (NFD) resources, or instance-oriented topology JSON. Cloud providers with a working integration include Google Cloud, Lambda, Nebius, Nscale, and OCI, with more cloud and colocation providers in development. On-premises support covers InfiniBand via ibnetdiscover, or NetQ for Spectrum-X or Multi-Node NVLink (MNNVL) domains. The provider interface is open for operators to add their own and contribute upstream.

Within the NVIDIA DSX OS cluster orchestration layer, Topograph works alongside Dynamic Resource Allocation (DRA) and KAI Scheduler to enable topology-aware gang scheduling. A Crusoe provider reads fabric and accelerator-domain labels from Crusoe Managed Kubernetes nodes, so Topograph runs in Kubernetes for that provider. The NFD engine requires the alpha NodeFeatureGroupAPI feature gate.

Five components keep the topology view current: an API Server, Node Observer, Node Data Broker, Provider, and Engine. The API server exposes endpoints including POST /v1/generate, GET /v1/topology, POST /v1/lookup, GET /healthz, and GET /metrics. A typical aggregation delay is 15 seconds, and repeated identical requests are processed once.

Deployment is available via Helm on Kubernetes, native packages on Slurm, and as a Helm chart for Slinky, developed by SchedMD; NVIDIA acquired SchedMD in December 2025. Supported Kubernetes prerequisites include version 1.27 or later and Helm 3.10+ or 4.x. Simulation utilities such as kwok-nodes and Kind/KWOK helpers allow testing without production hardware.

NVIDIA states that Topograph reflects reported rather than intended topology, and that labels refresh when generation runs. A Slurm strigger script registers on node up and down transitions but does not detect arbitrary switch rewiring or every inventory change. Topograph is available from the dsx-ai-factory/topograph GitHub repository.

Read at NVIDIA Developer Blog

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

AI factories are power-limited systems that deliver maximum value when fully optimized. GPU workload placement is a key optimization. Poor workload placement...