AivexaNewsSearch
AI news for builders and product teamsChecked every hour

Up to 3.2x Faster Inference with LFM2.5-DSpark

Collected Oct 1, 2026

Liquid AI has released DSpark draft model checkpoints for three models in its LFM2.5 family: LFM2.5-1.2B-Instruct, LFM2.5-2.6B, and LFM2.5-8B-A1B. The checkpoints add a speculative decoding path that the company says trades a minimal memory increase for a large decoding speedup without changing output quality.

According to Liquid AI, the draft models deliver up to 3.18x throughput improvement on a GPU and up to 2.87x on-device. For LFM2.5-2.6B, DSpark reduces function-calling latency by 57% on average across multi-tool scenarios. Support for llama.cpp and SGLang is open-sourced upstream from day one.

DSpark combines a DFlash-style parallel backbone conditioned on the target model's context features, a lightweight sequential head modeled as a Markov chain between neighboring tokens, and a confidence-scheduled verifier that prunes low-confidence suffixes. The draft models are attention-only with 5 layers and a block size of 9, each around 300M parameters. Liquid AI trained each for 15 epochs on a data mix covering SFT, chat, code, and function-calling data, selecting the epoch with the highest acceptance rate.

Under greedy decoding, a draft token is accepted only if it matches the target model's distribution, so the emitted sequence is identical to baseline greedy decoding and benchmark accuracy is unchanged, the company said.

Measurements used llama.cpp with Metal on an M4 Max MacBook Pro with FP16 GGUF weights and up to 256 output tokens, and SGLang on a single H100 80 GB in BF16, both at block size 9, batch size 1, temperature 0, across five benchmark datasets. Liquid AI reported roughly 140 tokens per second on the MacBook for LFM2.5-2.6B. For LFM2.5-1.2B-Instruct, speedup varied by as much as 52% depending on text distribution. For LFM2.5-8B-A1B, on-device improvement averaged 18%, which the company attributed to the current MoE implementation in llama.cpp's Metal backend and to verifying multiple tokens activating more experts.

Checkpoints are available on Hugging Face in Safetensors and GGUF formats.

Read at Hugging Face Blog

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt