AivexaNewsSearch
AI news for builders and product teamsChecked every hour

Faster Gemma 4 on MLX with multi-token prediction

Collected Oct 1, 2026

Ollama 0.31 adds multi-token prediction (MTP) support for Gemma 4 on Apple Silicon, powered by MLX, which the company says makes the model generate tokens nearly 90% faster on average across a coding-agent benchmark. The speedup is on by default and does not change the model's output, according to the announcement dated June 29, 2026.

MTP works by pairing the main model with a small, fast draft model that proposes the next several tokens. The main model verifies the proposal in a single pass and keeps the tokens it agrees with. Because the draft model is a small fraction of the main model's size, its proposals are inexpensive; when correct, several tokens are committed for the cost of one. Ollama notes code is especially predictable, containing closing brackets, repeated identifiers, and boilerplate, so proposals are accepted often.

Ollama says the hard part is doing this reliably, since the ideal number of tokens to draft varies moment to moment and drafting too many can make MTP slower than plain decoding. Ollama tunes draft length automatically at runtime, tracking acceptance rates and verification pass times, and falls back to one-at-a-time decoding when proposals stop being accepted.

The engine runs drafting, sampling, verification, and subsequent sampling on the GPU as a single pass, with no return to the CPU. Ollama says the engine records a rollback point before each proposal, so a rejection rewinds to the last accepted token without touching or recomputing earlier work.

Ollama says most of the cost is in verification rather than drafting, and that the verification batch, usually 2 to 8 tokens, falls between the single-token and large-batch sizes typical matrix multiplication kernels target. Ollama contributed a kernel to MLX that reads and unpacks each block of weights once and reuses it across the batch. On an M5 Max with nvfp4, Ollama says this makes Gemma 4's largest matrix multiplications 2x to 2.5x faster, with identical computation and the speedup coming from removing redundant work. The kernel is available to other models in MLX, not only Gemma 4 in Ollama.

The measurements were taken on the Aider polyglot benchmark. Ollama states the benefit depends heavily on the workload. The announcement says Gemma 4 is the first model to receive this performance improvement, with more to follow.

Read at Ollama

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

Gemma 4 is now significantly faster in Ollama 0.31 on Apple Silicon via multi-token prediction (MTP), powered by MLX. Performance is now up to 90% faster when used with coding agents, as measured using the Aider polyglot benchmark.