Introducing @huggingface/kernels: 200+ WebGPU Kernels for Local AI
Hugging Face has released @huggingface/kernels, a minimal JavaScript library for loading and running optimized WebGPU kernels from the Hugging Face Hub, alongside an initial collection of 207 kernels published in the webgpu-kernels organization. The kernels are Apache-2.0 licensed and each is published as its own repository and kernel card documenting the operation's semantics, inputs, outputs, attributes, supported data types, source files, and a ready-to-run example.
Each kernel repository packages a manifest.json defining the operation contract, metadata.json recording the kernel identifier, digests and provenance, test.json with correctness cases, bench.json with benchmark and tuning cases, and WGSL shader template files. The library is installed from npm via npm install @huggingface/kernels@preview and requires a browser with WebGPU support.
Hugging Face also launched Fleet, an in-browser GPU benchmarking and testing suite that runs and scores the kernels on a user's hardware. With user consent, each run contributes private evidence intended to help find failures such as incorrect results or slow cases, improve kernel variants, and inform optimization decisions.
In a comparison against ORT WebGPU on an Apple M4 GPU using ONNX Runtime Web 1.30.0-dev.20260826-b1f76d586a, Hugging Face started with 1,756 test cases across all 207 operations and kept 809 cases where both sides produced matching outputs and reliable timings. Across those, the kernels were 2.57x faster by geometric mean and 1.90x faster at the median, with 629 wins, 176 losses, and 4 ties. Two cited examples: a bilinear Einsum case ran in 0.136 ms versus 1,396 ms, and a row-wise CumSum over [256, 4096] was 0.016 ms versus 4.784 ms. The timings measured GPU work only, excluding setup such as loading kernels, creating sessions, uploading inputs, compiling shaders, and reading outputs back. Hugging Face said it is working with the ONNX Runtime team to upstream these improvements.
Based on reporting from the original publisher. Visit the source for full context and later updates.