Ollama is now powered by MLX on Apple Silicon in preview
Ollama is previewing a version powered by MLX, Apple's machine learning framework, on Apple Silicon. The company said Ollama on Apple Silicon is now built on MLX to take advantage of its unified memory architecture, resulting in what it described as a large speedup across all Apple Silicon devices.
On Apple's M5, M5 Pro and M5 Max chips, Ollama uses the new GPU Neural Accelerators to accelerate both time to first token and generation speed, measured in tokens per second. Testing was conducted on March 29, 2026, using Alibaba's Qwen3.5-35B-A3B model quantized to NVFP4 and Ollama's previous implementation quantized to Q4_K_M using Ollama 0.18. The company said Ollama 0.19 will see even higher performance, reporting 1851 tokens/s prefill and 134 tokens/s decode with int4 quantization.
Ollama now uses NVIDIA's NVFP4 format to maintain model accuracy while reducing memory bandwidth and storage requirements for inference workloads. Ollama said this lets users share the same results as in a production environment, and opens the ability to run models optimized by NVIDIA's model optimizer. Other precisions will be made available based on design and usage intent from Ollama's research and hardware partners.
The cache was upgraded for coding and agentic tasks: Ollama reuses its cache across conversations, stores snapshots at intelligent locations in the prompt, and preserves shared prefixes longer when older branches are dropped. Ollama said this reduces memory utilization and prompt processing.
The preview release accelerates the new Qwen3.5-35B-A3B model with sampling parameters tuned for coding tasks, and requires a Mac with more than 32GB of unified memory. Ollama said it is working to support future models, will introduce an easier way to import custom models fine-tuned on supported architectures, and will expand its list of supported architectures. The company credited the MLX contributor team, NVIDIA contributors, the GGML and llama.cpp team, and the Alibaba Qwen team.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
Today, we're previewing the fastest way to run Ollama on Apple silicon, powered by MLX, Apple's machine learning framework.