AivexaNewsSearch
AI news for builders and product teamsChecked every hour

Building Spyre as a Native PyTorch Device

Collected Oct 8, 2026

IBM's torch-spyre team has integrated Spyre, IBM's dataflow AI accelerator for inference, as a native PyTorch device. The work maps PyTorch's device, allocator, stream, event and launch abstractions onto the Spyre runtime and firmware, giving Spyre a real device identity through PrivateUse1 registration. Spyre targets enterprise teams running AI alongside applications and data on IBM Z, LinuxONE and Power systems, with reduced-precision compute suited to matrix-heavy language generation and embedding workloads.

The hardware has 32 cores connected by a high-bandwidth ring, each with 2 MB of local scratchpad, and up to 128 GB of LPDDR5 for tensors and programs. The runtime stack owns LPDDR5, while the compiler emits the loads and stores that move tiles between LPDDR5 and each core's scratchpad in 128-byte sticks. Two registration calls, torch.utils.rename_privateuse1_backend("spyre") and torch._register_device_module("spyre", make_spyre_module()), enable tensor.to("spyre") and dispatcher routing.

Spyre's constraints shape the design. Computation is data-triggered: a compiled kernel contains programs for multiple functional units, and the compiler builds the device-side producer/consumer schedule. Independently submitted compute operations share one runtime compute queue, so overlap comes from running data movement on separate pipelines rather than concurrent compute launches; transfer and compute work go on different streams, joined by events where one consumes the other's result. Work is submitted as ordered queues of typed operations.

Memory is managed in regions, contiguous chunks identified by a handle. A card can be dedicated to one tenant (PF mode) or shared (VF mode, with the tighter handle budget). The allocator acquires a few large regions and sub-allocates aligned blocks, so many tensors share few handles. An allocation is a region identifier, offset and length; tensors interleaved across memory domains comprise several such pieces, carried in an opaque context alongside the pointer that PyTorch reference counts.

Programs are compiled ahead of time against fixed layouts, but caller-owned tensor addresses stay symbolic until the allocator places them. The current runtime patches the program at launch: a host callback writes resolved addresses into a pinned host buffer, a transfer copies it into the program's allocation, and the computation patches its operands. torch.compile(backend="inductor") keeps FX graphs in Inductor; a compiled artifact is prepared once into a per-artifact launch plan, so launch is only operation construction and enqueue.

Why it matters: Developers get one path for eager and compiled execution on Spyre, with lower launch overhead and tensors that stay resident on device between operations. That removes per-operation transfer boundaries for eager work and avoids a second runtime graph that previously left PyTorch unaware of device residency.

Read at PyTorch

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

TL;DR Spyre becomes a native PyTorch device by connecting PyTorch’s existing device, allocator, stream, and compiler abstractions through torch-spyre to the Spyre runtime and firmware. PrivateUse1 gives Spyre a real... The post Building Spyre as a Native PyTorch Device appeared first on PyTorch .