AivexaNewsSearch
AI news for builders and product teamsChecked every hour

Improved performance and model support with GGUF

Collected Oct 1, 2026

Ollama 0.30 is now available, adding improved performance and GGUF model compatibility through llama.cpp. The release augments Ollama's MLX engine on Apple silicon and is intended to bring support to more models across a wider range of hardware.

On NVIDIA hardware, Ollama says performance is now up to 20% faster, using optimizations contributed by the NVIDIA and llama.cpp teams. The company said this was tested with the Gemma 4 26B model running on an NVIDIA RTX 5090 using the Q4_K_M quantization.

Vulkan is now enabled by default, which Ollama says extends its GPU acceleration to a wider range of hardware, including AMD and Intel devices. According to the company, more users can run models on the GPU out of the box without installing vendor-specific libraries.

Ollama 0.30 also expands compatibility with the GGUF ecosystem, so more models run out of the box, including model families such as LFM and Prism, as well as fine-tuned models published by Unsloth.

To run a GGUF model from Hugging Face, Ollama instructs users to first download the GGUF file or a directory containing GGUF files, then create a Modelfile with the FROM command pointing to the path of the GGUF file or directory. Users then create and run the model with the ollama create and ollama run commands.

If a model supports tool calling, that capability carries over to Ollama, according to the company, and such models can be used with coding agents and personal assistants in a single command, including Claude Code, Hermes Agent and OpenClaw via ollama launch. Ollama says the tools capability can be verified with ollama show.

Ollama acknowledged the work of Georgi Gerganov and the llama.cpp maintainer teams, as well as hardware partners including NVIDIA, AMD, Qualcomm and Intel.

Read at Ollama

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

Ollama 0.30 is now available with improved performance and GGUF model compatibility through llama.cpp. This augments Ollama's MLX engine on Apple silicon, bringing support to more models on a wider range of hardware.