Open Research, Tooling & Optimization at PyTorch Conference North America 2026
PyTorch Conference North America 2026 is scheduled for October 20–21 in San Jose, California. The event's program centers on open research, tooling, and optimization, spanning the Core PyTorch, Kernel Engineering, Training, and Inference tracks, according to the conference announcement.
The compiler-focused sessions cover torch.compile, Dynamo, and dynamic shapes. Scheduled talks include a device-aware tensor layout extension for Inductor (Olivier Tardieu and Matthew Arnold, IBM), new Dynamo support for nested graph breaks (William Wen, Meta), static tensor shape checking with Pyrefly evaluated across 28 real LLM, vision, recommender, and RL models (Steven Troxler and Avik Chaudhuri, Meta), parametrized dynamic shape CUDA graphs (Elias Ellison, Meta; Daniel Galvez, NVIDIA), and a new C++ FakeTensor implementation reported to deliver a 30x speedup on operations like aten.mm (Angel Li, Meta).
Kernel engineering and domain-specific language sessions feature AMD's FlyDSL integration into TorchInductor's GEMM pipeline (Liz Li), Helion backends for CuteDSL on recent NVIDIA GPUs and Pallas for TPUs (Oguz Ulgen, Dunfan Lu, Jason Ansel, Meta), Helion autotuning advances including Likelihood-Free Bayesian Optimization reported to cut tuning time by 36.5% (Jongsok Choi, Ethan Che, Meta), KernelAgent's hardware-guided multi-agent workflow reporting 1.56x speedup over default torch.compile and 89% of H100 roofline efficiency (Kaiming Cheng, Laura Wang, Meta), and Hugging Face's Kernels library (Sayak Paul).
Distributed training and communication sessions include Monarch (Marius Eriksen, Meta), FlexShard (Wei Feng, Anshul Sinha, Ailing Zhang, Meta), DeepSpeed's AutoTP, AutoSP, and AutoEP (Masahiro Tanaka, Anyscale), Precompile for Training (Bob Ren, Aaron Orenstein, Meta), AutoParallel (Sanket Jayant Purandare, Francisco Massa, Meta), EP-Overlap (Sanket Jayant Purandare, Meta), rocSHMEM symmetric memory for AMD GPUs (Prachi Gupta, AMD), XCCL on Intel GPUs validated on Argonne's Aurora (Panagiotis Kourdis, Tanima Dey, Intel), NCCL Extensions (Sreeram Potluri, Artem Polyakov, NVIDIA), and MCCL (Ben Carver, Meta).
Quantization sessions include HiFloat8 and HiFloat4 (Yun Zhao, Haonan Zhang, Huawei) and NVFP4 pretraining recipes upstreamed into TorchAO and TorchTitan (Anjulie Agrusa, Ryan Spring, Bruce Zitelli, NVIDIA). Other sessions cover optimizers beyond AdamW, including Muon, Dion, and orthogonalized variants.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
TL;DR Taking place October 20 to 21 in San Jose, California, PyTorch Conference North America 2026 highlights open research, tooling, and performance optimization across compiler architecture, cross-hardware kernel domain-specific languages,...