AI news for builders and product teamsUpdated Sep 30, 2026, 20:01 UTC
NVIDIA Developer Blog
First-party releases and research from NVIDIA Developer Blog. Headlines and excerpts link to the original articles.
Latest stories
Newest firstND
NVIDIA announced general availability of its cuObject client and server libraries and expanded the xio-sig consortium to include cuObject alongside cuFile, with Google Cloud evaluating participation and Microsoft planning to join the board. A new SCADA Server SDK lets storage providers build servers for GPU-initiated requests, and IBM demonstrated a prototype integrating SCADA with IBM Storage Scale.
ND
An agent can finish a task and still take an inefficient path. A failed search can trigger another search. A truncated file read can lead to a command fetching...
ND
Parallel work, model-family isolation, reversible changes, and GPU-backed validation shaped an open source project designed around coding agents NVIDIA TensorRT...
ND
Vision-language models have made it possible to build visual AI agents that understand video at production scale. The harder problem is turning that capability...
ND
NVIDIA introduced its Open Agent Safety Platform, combining the open-source OpenShell secure runtime with NVIDIA Sentry on BlueField-4 DPUs to monitor and enforce policies for AI agents. It is built around five stated principles and is optimized for NVIDIA Vera CPU and BlueField DPU systems.
ND
NVIDIA released OpenShell 0.1.0, an open-source runtime that enforces which systems and data an AI agent can access without rewriting the agent. It combines sandboxed execution, controlled service access, credential management and formal policy analysis; Cadence, Slack and Gecko Robotics are adopting it.
ND
NVIDIA described a joint evaluation with Nscale of its DSX MaxLPS policy-governed power sharing on GB300 NVL72 systems running Kimi K2.5 workloads, reporting 49.2% higher normalized aggregate throughput and up to 40% more GPUs within the same approved power budget, with P99 time to first token up 17%.
ND
NVIDIA published a tutorial on training mixture-of-experts biological foundation models using Transformer Engine primitives, including GroupedLinear and a fused MXFP8 GroupedMLP kernel. A benchmark on eight NVIDIA B200 GPUs reported up to 2.21x the throughput of a Hugging Face baseline.
ND
NVIDIA released NV-Reason-CT, an open 3D CT vision language model that generates radiologist-style chain-of-thought reasoning. It combines a 3D ViT encoder with a Qwen3.5-4B language model and reports state-of-the-art results on the CT-RATE benchmark, with NIH radiologists validating the clinical plausibility of its reasoning traces.
ND
NVIDIA Cluster Readiness Engine (NVCRE) is an open source Kubernetes controller that runs real distributed workloads to validate GPU cluster readiness before production workloads. It attributes failures to specific nodes and categories, supports adaptive fault isolation, and integrates with NVIDIA AI Cluster Runtime and NVSentinel.
ND
NVIDIA introduced NodeWright, an open-source, Kubernetes-native package manager that declaratively configures and updates host operating systems across GPU clusters. It orchestrates cordon, drain, apply, interrupt, and uncordon sequences per node and supports progressive fleet rollouts through DeploymentPolicy resources.
ND
NVIDIA introduced SWE-Serve, a benchmark of 53 tasks derived from 83 merged SGLang pull requests, developed with input from the SGLang team. On 19 tasks with live-serving checks, the same patches passed 45.9% with the complete verifier versus 69.4% without those checks.
ND
NVIDIA described Confidential Computing adaptations in TensorRT LLM for private AI inference on Blackwell GPUs. On eight B200 GPUs with DeepSeek-R1, confidential compute retained 96.1-98.2% of baseline output-token throughput with 1.2-4.3% per-token latency overhead across concurrency 1-16.
ND
NVIDIA announced Topograph, an open source toolkit that discovers cluster network topology and publishes it in formats schedulers can consume, including Kubernetes node labels, Slurm configuration, and Slinky ConfigMaps. It supports cloud providers Google Cloud, Lambda, Nebius, Nscale, and OCI, plus on-premises InfiniBand, Spectrum-X, and Multi-Node NVLink domains.
ND
NVIDIA DLSS 5 introduces DLSS 3D-Guided Neural Rendering and granular controls that help game developers add lifelike lighting and material detail while...
ND
GPU acceleration can speed up compute-intensive robotics workloads, but a fast CUDA kernel alone does not guarantee a fast ROS 2 graph. As messages move between...
ND
NVIDIA Dynamo-Triton release 26.07 enables the TensorRT backend multi-device capability, letting one KIND_MODEL instance own multiple GPUs and serve distributed inference through a single gRPC endpoint. A Cosmos 3 Nano demonstration cut end-to-end generation latency from 156.6 seconds on one GPU to 34.2 seconds on eight GPUs.
ND
NVIDIA published a developer blog explaining how AI agent evaluation has shifted from scoring single function calls to measuring full task completion in executable environments. It describes step-level and end-to-end scoring on execution traces, a benchmark-trial-task-turn-step metric hierarchy, and reports Nemotron 3.5 Lightning at 86% accuracy on PinchBench while finishing tasks 30% faster than Qwen3.6 35B at comparable accuracy.
ND
NVIDIA published a tutorial on AI data assimilation tools in its Earth-2 platform, covering Score-Based Data Assimilation for regional diffusion models and HealDA for global atmospheric state estimation. It reports wind-speed RMSE reductions of 54% in one CorrDiff-COSMO downscaling example and an average of 7.2% across six StormCast-CONUS forecast steps.
ND
NVIDIA introduced AIPerf, a multiprocess LLM inference benchmarking tool that succeeds GenAI-Perf and avoids client-side bottlenecks under high concurrency. It supports 15+ endpoint types, public datasets and trace replay, configurable arrival patterns, and reports TTFT, ITL, request latency and output token throughput with percentile breakdowns and optional GPU telemetry.