AI news for builders and product teamsUpdated Sep 30, 2026, 20:01 UTC
NVIDIA Developer Blog
First-party releases and research from NVIDIA Developer Blog. Headlines and excerpts link to the original articles.
Latest stories
Newest firstND
NVIDIA describes an agentic AI workflow that prepares Blender 3D scenes for robotics simulation, using Codex or Claude to coordinate subagents that add semantic labels, physics properties and sensors, then hand off OpenUSD scenes to Isaac Sim or Isaac Lab.
ND
NVIDIA said TensorRT Edge-LLM ran Qwen3.6-27B on a single Jetson AGX Thor Developer Kit in the MLPerf Inference v6.1 Edge Agentic benchmark at 52.33 tokens per second, completing 1,007 turns in 24 minutes 36 seconds. NVIDIA reported this as 6.4x faster than the llama.cpp reference run of 2 hours 37 minutes.
ND
NVIDIA describes an agent skill in its TileGym repository that translates cuTile Python and Triton-TileIR GPU kernels into cuTile Rust. All 24 public TileGym operators were ported, reaching 99.5% of cuTile Python performance on average on NVIDIA DGX B200 hardware.
ND
How can a 30B-parameter model activate only 3B parameters per token, and still use the capacity of the larger model? Nemotron 3.5 Lightning illustrates the...
ND
For operators of large-scale AI factories, maximizing continuous output is essential for productivity. In massive-scale AI training, every GPU in the cluster...
ND
Power is a defining constraint for AI factories. As AI workloads demand a full compute platform to serve them, each component of that platform must maximize...
ND
Federated learning (FL) projects often begin with a straightforward setup: one server, a few clients, and one dataset at each site. As those projects grow, the...
ND
NVIDIA reports that its Transformer Engine with JAX raises DeepSeek-V3 MoE training throughput from 103 to 1,068 TFLOPS/GPU on GB200, a 10.4x gain. The stack sustains 97% scaling efficiency at 1,024 GPUs on GB300 NVL72.
ND
NVIDIA says its NIM 2.0.12 optimized serving stack delivers up to 2.5x higher system throughput on a 4xB200 system versus a baseline without NIM optimizations, reaching 1,997 tokens per second at a 50 TPS per user target for Nemotron 3 Ultra.
ND
Biomolecular structure prediction is now often run at proteome scale, where the goal is to move an entire worklist through the pipeline efficiently. NVIDIA...
ND
NVIDIA and Palantir built a Digital Supply Chain Intelligence command center on Palantir Foundry, using NVIDIA cuOpt for weekly allocation optimization and post-training a 30B Nemotron 3.5 Lightning model on captured planner decisions. On a development benchmark, the post-trained model reached 86.7% allocation-decision accuracy.
ND
NVIDIA's developer blog describes encode-prefill-decode (EPD) disaggregation in NVIDIA Dynamo, which separates vision encoding from LLM prefill and decode for multimodal serving, reporting up to 5x faster time to first token and 7x faster end-to-end response time in tested scenarios.