AivexaNewsSearch
AI news for builders and product teamsChecked every hour

ParallelKernelBench: Frontier LLMs can't write fast multi-GPU kernels (yet)

Collected Oct 1, 2026

Together AI has introduced ParallelKernelBench (PKB), a benchmark and evaluation framework for multi-GPU kernel generation. PKB includes 87 problems drawn from real codebases such as Megatron-LM, DeepSpeed, DeepEP, TensorRT-LLM, and NeMo-RL, where the task is to replace a PyTorch + NCCL implementation with a CUDA kernel that moves data directly over NVLink using symmetric memory.

Frontier coding models tested include GPT-5.5, Gemini 3 Pro, and Opus 4.7. In the zero-shot setting, the best model solved 28 of 87 problems correctly, and only 22 of those solutions were faster than the PyTorch + NCCL baseline. With three attempts sampled, the best result improved to 36 correct solutions and 27 faster-than-baseline solutions, but fast1@3 still peaked at 31%.

Successes concentrated in collective primitives, tensor-parallel GEMMs, and Ulysses-style context parallelism. Weaker models often failed to compile, while stronger reasoning models frequently produced kernels that compiled but returned incorrect results. Generated kernels mostly relied on copy engines or SM load/store instructions, with mechanisms such as TMA and NVLS almost absent.

An agentic harness wrapping Gemini 3 Pro improved results from 24 to 35 correct solutions out of 87, with 26 kernels beating the baseline, but performance plateaued after roughly 20 refinement steps.

Single-shot generation occasionally produced kernels faster than any publicly available implementation, including one for NVIDIA NeMo-RL's GRPO training loop that had no prior optimized public reference. Three examples are highlighted: a NeMo vocab-parallel log-prob kernel with top-k/top-p filtering (Gemini 3 Pro), a Hyena forward context parallelism kernel (GPT-5.5), and a SAM 3 all-gathered mask IoU suppression kernel (GPT-5.5). Each was verified for correctness over 4 H100 GPUs and 100 randomized runs.

PKB is scoped to intra-node NVLink today, with planned extensions to inter-node fabrics and other accelerators. The benchmark is being released as open, and contributions, especially inter-node problems, are invited.

Read at Together AI

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

ParallelKernelBench tests whether LLMs can write fast multi-GPU CUDA kernels across 87 real workloads. The best model solves under a third, but a few generated kernels beat any public implementation.