Building a High-Performance and Portable vLLM Linear Backend with Helion

PyTorch developers added a Helion linear backend to vLLM, the LLM inference and serving framework, to test whether an autotuned, high-level kernel DSL can improve inference performance while reducing kernel implementation complexity. The work targets NVIDIA Hopper GPUs through Helion's Triton backend and covers the FP8_Dynamic, W8A8_INT8 and Block_FP8 quantized GEMM formats. On kernel-level benchmarks, it reports geometric mean speedups of 1.110x for FP8_Dynamic over CUTLASS, 1.178x for W8A8_INT8 over CUTLASS, 1.149x for Block_FP8 over FlashInfer, and 1.177x for Block_FP8 over DeepGEMM.
The core idea is that a single Helion GEMM implementation covers three algorithmic variants: Standard GEMM, Split-K, which partitions the K dimension across thread blocks when the M or N dimensions are too small for parallelism, and Swap-AB, which rewrites A@B as (B.T@A.T).T for small-M shapes. The variants are exposed as tunable parameters instead of separate kernels and hand-written dispatch heuristics. In the shown matmul kernel, split_k is constrained to power-of-two values up to 256 and swap_ab is Boolean. An ahead-of-time autotuner benchmarks combinations and selects the best per shape. Initial results with Helion's CuteDSL backend show competitive GEMM performance on NVIDIA Blackwell GPUs, and the work could extend there as that backend matures.
The backend uses hybrid dispatch: for small shapes up to max_helion_size, set to 32 for this work, it dispatches Helion under CUDA Graph replay; larger shapes fall back to CUTLASS or DeepGEMM. Helion kernels are autotuned for num_tokens of 1, 2, 4, 8, 16, 24 and 32, using the LLMSeededLFBOTreeSearch autotuner seeded by Claude Opus 4.8 with HELION_BENCHMARK_CUDAGRAPH=1. End-to-end benchmarks on an NVIDIA H100 80GB HBM3 GPU with ShareGPT used the vllm serve command with --max-num-seqs 32, --no-enable-prefix-caching and --linear-backend helion across Qwen3-1.7B through Qwen3-32B plus Qwen3.8-27B, showing consistent gains and more than 10% throughput improvement for some workloads.
Acknowledged challenges include hours-long ahead-of-time tuning, startup JIT compilation during CUDA Graph capture that warm-start caching can largely eliminate, CPU dispatch overhead outside CUDA Graphs, and maintenance of large pre-tuned config files.
Why it matters: Developers and serving teams can now tune kernels for their own deployments via a vLLM fork with the backend, autotuning tooling and instructions, while default backend users are unaffected. The team is exploring upstreaming the kernels and framework with a default config, leaving workload-specific autotuning to end users, a model it says favors performance and maintainability over out-of-the-box usability.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
TL;DR We integrated Helion into vLLM’s linear backend to explore how an autotuned, high-level kernel DSL can improve LLM inference performance while reducing kernel implementation complexity. A single Helion general... The post Building a High-Performance and Portable vLLM Linear Backend with Helion appeared first on PyTorch .