AI news for builders and product teamsUpdated Oct 10, 2026, 18:01 UTC
PyTorch
First-party releases and research from PyTorch. Headlines and excerpts link to the original articles.
Latest stories
Newest first
IBM's torch-spyre team has made Spyre a native PyTorch device via PrivateUse1, giving tensors a real device identity on device="spyre" backed by PyTorch's allocator, streams and events. Compiled artifacts now launch as prepared recipes of typed operations instead of a second runtime graph.

PyTorch's FBTriton reimplements the Table Batched Embedding (TBE) forward and backward passes in Triton, reporting a median forward speedup of 1.28x over legacy CUDA kernels across 307 GB200 shard configurations. On one large B200 configuration, combined latency dropped from 79.537 ms to 66.183 ms.

PyTorch has consolidated all media decoding and encoding across images, video, and audio into TorchCodec, deprecating or removing the equivalent APIs in TorchVision and TorchAudio, which now focus on transforms. All three libraries are now ABI stable and no longer tied to a single PyTorch version.

PyTorch's Accelerator Integration Working Group detailed H1 2026 progress, including the Cross-Repository CI Relay, migration of 276-plus test files to device-agnostic form, and OpenReg-based reference work for profiling, distributed and compile integration.

A PyTorch team integrated the Helion kernel DSL into vLLM's linear backend, letting one autotuned GEMM cover Standard, Split-K and Swap-AB variants; on NVIDIA Hopper it beats CUTLASS, DeepGEMM and FlashInfer kernels, with over 10% end-to-end throughput gains on some workloads.

PyTorch launched a new PyTorch Certified Associate (PTCA) Certification Pathway through Linux Foundation Education, bundling four self-paced learning modules with the PTCA certification exam in one structured product.

Meta's TLX-based Jagged Flash Attention kernel for GEM on NVIDIA B200 uses about 3.2K lines of Triton code, roughly 3x less than FA4's ~10K-line CuteDSL kernels, while beating FA4 by about 13% on the forward pass and 50% on the backward for jagged shapes.

PyTorch Conference North America 2026 in San Jose will feature a series of Ray-focused sessions covering Ray's integration with Kubernetes and production use cases at LinkedIn, Uber, Pinterest, and Anyscale.

Torch Spyre reached CRCR L2 integration by testing the PyTorch backend for IBM's Spyre accelerator against PyTorch's test suite, using an agentic test-selection pipeline and a declarative YAML framework that adapts tests without patching upstream. The design targets any generic privateuse1 device, and the authors intend to work with the PyTorch team to upstream reusable parts.

The PyTorch Foundation is adding a new Introduction Track to PyTorch Conference North America 2026 and hosting a full-day, instructor-led PyTorch Associate Training on Monday, October 19, 2026, in San Jose, California.

The PyTorch Foundation OSPO & Academic Outreach Working Group is inviting academic PyTorch projects to submit for a Day 0 Academic Workshop ahead of PyTorch Conference North America 2026. Eight selected projects will give five-minute lightning talks, with submissions due September 30, 2026, midnight PT.

vLLM is introducing new "HW agnostic" layers to keep the project portable across diverse hardware as its frontier optimizations move away from fullgraph torch.compile. On NVIDIA H100 GPUs, these layers achieve total token throughput within 3.4% of the native implementation (geometric mean across three recent models).

A PyTorch case study describes how Shopify built a continual learning loop using PyTorch and vLLM, turning production failures into model weight updates. Shopify reports its GraphQL agent surpassed frontier-model quality while cutting serving costs by 96% and reducing latency.

TinyTorch is a free, open-source curriculum from the PyTorch ecosystem where learners build a working ML framework in pure Python, tensors through transformers, across twenty modules on a laptop with 4 GB of RAM and no GPU. It launched in December 2025 and is in preview, aimed at classroom readiness for Fall 2026.

PyTorch Day Japan 2026 will be held in Tokyo on December 10, hosted by the PyTorch Foundation, Hugging Face, IBM, and Mitsubishi Electric. The call for proposals closes September 27, and discounted registration at ¥5,000 is available through November 11.

PyTorch Conference North America 2026 will take place October 20–21 in San Jose, California, with sessions on open research, tooling, and performance optimization across compiler architecture, cross-hardware kernel DSLs, exascale distributed training, and low-precision quantization.

PyTorch published a technical report on extending FlashAttention-4 with end-to-end MXFP8 block-scaled attention for Blackwell GPUs, reporting up to 1.6x forward and 1.52x backward gains over BF16. The code is open sourced in Meta's ads_model_kernel_library repository.