AivexaNewsSearch
AI news for builders and product teamsChecked every hour

Hardware-Agnostic Models in vLLM

Collected Oct 1, 2026

vLLM announced it is introducing a set of "HW agnostic" layers in-tree, aimed at preserving support for diverse models across diverse hardware as its internal implementation changes in ways incompatible with fullgraph torch.compile. The project said that to reach state-of-the-art frontier performance, existing layers and ops are likely to be refactored in a way that makes them fundamentally incompatible with torch.compile and removes extensibility via CustomOp.

vLLM currently offers three flavours of model definitions: new "flat" models under vllm/models/, legacy models under vllm/model_executor/models, and the transformers modeling backend. The project said it is starting to maintain hardware-specific "flat" model definitions that use custom fusions and model- and hardware-specific optimizations rather than torch.compile, and that new frontier models added in recent months all use this flat definition.

According to vLLM, this creates concerns for out-of-tree (OOT) plugins, older or more exotic models served via the transformers backend, and older or consumer/prosumer GPUs. Legacy model definitions are being removed from model_executor/models and registry entries updated to point at the transformers modeling backend.

The new HW agnostic layers follow four design principles: compilable, extensible via CustomOp and PluggableLayer, isolated from hardware-specific layers and ops, and portable using native PyTorch code or portable DSLs like Triton and Helion. For the transformers backend, vLLM said it modified its "rewiring" process to target the new layers at model_executor/hw_agnostic instead of model_executor/layers. Support has landed on main for a limited number of layers and can be enabled with USE_HW_AGNOSTIC=1, for example: USE_HW_AGNOSTIC=1 vllm serve google/gemma-4-31B --model-impl=transformers.

vLLM said it validated the pathway using the Spyre OOT plugin for models including Gemma 4, Qwen3, and Granite 4.2, and plans to include HW agnostic models in CI and gradually switch to using this as the default pathway for serving models on Spyre. It also plans a new model.py for each flat model using the HW-agnostic layers, with shared layers in model_executor/hw_agnostic and model-specific layers such as DeepSeekV4FlashMLAAttention in the model's local directory. A hardware-agnostic DeepSeek V4 PR is under review.

On NVIDIA H100 GPUs, vLLM said the HW agnostic layers achieve total token throughput within 3.4% of the native implementation (geometric mean across three recent models). It stressed that state-of-the-art performance on Blackwell, CDNA 4, and beyond is not the goal of these model definitions; the aim is platform and performance portability across diverse hardware, including OOT accelerators, older GPUs, and prosumer-grade GPUs. The effort is described as a work-in-progress, with feedback sought via an RFC and the #hw-agnostic-models channel on vLLM Slack.

Read at PyTorch

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

TL;DR To achieve state-of-the-art performance at the frontier, vLLM is changing its internal implementation in ways that make it incompatible with fullgraph torch.compile. This may have consequences for users who...