AivexaNewsSearch
AI news for builders and product teamsUpdated Sep 30, 2026, 20:01 UTC

NVIDIA Developer Blog

First-party releases and research from NVIDIA Developer Blog. Headlines and excerpts link to the original articles.

Latest stories

Newest first
NVIDIA Developer BlogFirst partyDeveloper tools

Expanding AI Storage Access with NVIDIA cuObject and the NVIDIA SCADA Server SDK

NVIDIA announced general availability of its cuObject client and server libraries and expanded the xio-sig consortium to include cuObject alongside cuFile, with Google Cloud evaluating participation and Microsoft planning to join the board. A new SCADA Server SDK lets storage providers build servers for GPU-initiated requests, and IBM demonstrated a prototype integrating SCADA with IBM Storage Scale.

Read original
NVIDIA Developer BlogFirst partyDeveloper tools

Validate GPU Cluster Readiness Before AI Workloads Land

NVIDIA Cluster Readiness Engine (NVCRE) is an open source Kubernetes controller that runs real distributed workloads to validate GPU cluster readiness before production workloads. It attributes failures to specific nodes and categories, supports adaptive fault isolation, and integrates with NVIDIA AI Cluster Runtime and NVSentinel.

Read original
NVIDIA Developer BlogFirst partyDeveloper tools

Manage Kubernetes Node Fleets with NodeWright

NVIDIA introduced NodeWright, an open-source, Kubernetes-native package manager that declaratively configures and updates host operating systems across GPU clusters. It orchestrates cordon, drain, apply, interrupt, and uncordon sequences per node and supports progressive fleet rollouts through DeploymentPolicy resources.

Read original
NVIDIA Developer BlogFirst partyDeveloper tools

Topology-Aware Workload Scheduling with NVIDIA Topograph

NVIDIA announced Topograph, an open source toolkit that discovers cluster network topology and publishes it in formats schedulers can consume, including Kubernetes node labels, Slurm configuration, and Slinky ConfigMaps. It supports cloud providers Google Cloud, Lambda, Nebius, Nscale, and OCI, plus on-premises InfiniBand, Spectrum-X, and Multi-Node NVLink domains.

Read original
NVIDIA Developer BlogFirst partyDeveloper tools

Simplifying Model Serving Across Multiple GPUs with NVIDIA TensorRT Multi-Device Integration in NVIDIA Dynamo-Triton

NVIDIA Dynamo-Triton release 26.07 enables the TensorRT backend multi-device capability, letting one KIND_MODEL instance own multiple GPUs and serve distributed inference through a single gRPC endpoint. A Cosmos 3 Nano demonstration cut end-to-end generation latency from 156.6 seconds on one GPU to 34.2 seconds on eight GPUs.

Read original
NVIDIA Developer BlogFirst partyDeveloper tools

How to Evaluate AI Agents From Tool Calls to Task Completion

NVIDIA published a developer blog explaining how AI agent evaluation has shifted from scoring single function calls to measuring full task completion in executable environments. It describes step-level and end-to-end scoring on execution traces, a benchmark-trial-task-turn-step metric hierarchy, and reports Nemotron 3.5 Lightning at 86% accuracy on PinchBench while finishing tasks 30% faster than Qwen3.6 35B at comparable accuracy.

Read original
NVIDIA Developer BlogFirst partyDeveloper tools

Turn Your Latest Observations Into Timely Weather Decisions With NVIDIA Earth-2

NVIDIA published a tutorial on AI data assimilation tools in its Earth-2 platform, covering Score-Based Data Assimilation for regional diffusion models and HealDA for global atmospheric state estimation. It reports wind-speed RMSE reductions of 54% in one CorrDiff-COSMO downscaling example and an average of 7.2% across six StormCast-CONUS forecast steps.

Read original
NVIDIA Developer BlogFirst partyDeveloper tools

Benchmarking LLM Inference at Scale with AIPerf

NVIDIA introduced AIPerf, a multiprocess LLM inference benchmarking tool that succeeds GenAI-Perf and avoids client-side bottlenecks under high concurrency. It supports 15+ endpoint types, public datasets and trace replay, configurable arrival patterns, and reports TTFT, ITL, request latency and output token throughput with percentile breakdowns and optional GPU telemetry.

Read original