Ollama's highest performance on Apple Silicon yet with MLX
Ollama announced an update to its MLX engine that it describes as its highest performance on Apple Silicon yet. According to Ollama, the engine leans more heavily on Apple's unified memory and the Metal-backed MLX framework, and models now produce higher quality responses, respond faster, and use less memory.
The update adds support for NVIDIA's model-optimized NVFP4 format. Ollama says this allows higher quality outputs than other 4-bit quantization formats while maintaining state-of-the-art performance, and lets models optimized for datacenter deployment be imported and run with the MLX engine, providing portability between datacenter and desktop. Ollama measured perplexity differences among q4_K_M, NVFP4, and unquantized bf16 weights for the Gemma 4 12B model, stating that NVFP4 roughly halves the quality loss of 4-bit quantization relative to unquantized BF16. NVFP4 is described as tracking the local dynamic range of model weights more closely, reducing quantization loss.
Ollama says the engine is up to 20% faster due to new optimizations, including fusing several operations into single Metal kernels through MLX's just-in-time compiler features and reworking GPU-backed sampling. It reports NVFP4 generates about 20% faster than q4_K_M on the updated engine, based on average output speed over 10 runs with an 8,300-token input prompt.
A new snapshot system saves model state at key points across conversations, using the same approach Ollama's cloud uses for agent workloads. Ollama describes scenarios including multiple agents resuming from their own saved states, thinking models resuming after reasoning tokens are dropped, and branching or retries diverging from cached conversations. The company notes that sliding-window attention and recurrent layers carry state that cannot be rewound, and that snapshots are saved where conversations are likely to return: at branches, at intervals through long prompts, and just before each response.
To run models with the MLX engine, Ollama directs users to download the latest version and run ollama run gemma4:12b-mlx, or use ollama launch for a coding agent.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
Ollama's MLX engine has been updated to deliver its highest performance on Apple Silicon yet. Models output higher quality responses, respond faster, and use less memory.