New model scheduling
Ollama announced a new model scheduling system on September 23, 2025. According to the company, the new engine measures the exact amount of memory required before running a model, rather than relying on an estimation as previous versions did.
Ollama states this exact memory management means over-allocations no longer occur, significantly reducing crashes caused by out-of-memory issues. It also says the new management allocates more memory to the GPU, increasing token generation and processing speeds, and schedules models more efficiently across multiple GPUs, improving multi-GPU and mismatched GPU performance. Ollama further says measurements in tools like nvidia-smi will now match ollama ps, making memory utilization easier to track.
Ollama published benchmark examples. For long context on one NVIDIA GeForce RTX 4090 with gemma3:12b at 128k context length, token generation rose from 52.02 to 85.54 tokens/s, VRAM from 19.9GiB to 21.4GiB, and layers loaded on GPU from 48/49 to 49/49. For image input on two NVIDIA GeForce RTX 4090 GPUs with mistral-small3.2 at 32k context length, prompt evaluation rose from 127.84 to 1380.24 tokens/s, token generation from 43.15 to 55.61 tokens/s, VRAM from 19.9GiB to 21.4GiB, and layers loaded on GPU from 40/41 to 41/41 plus the vision model.
Ollama says all models implemented in its new engine have the feature enabled by default, with more models coming soon as they transition to the new engine. Supported models listed include gpt-oss; llama4 and llama3.2-vision, with llama3.2, llama3.1 and llama3 soon; gemma3, embeddinggemma and gemma3n; qwen3 and qwen2.5vl, with qwen3-coder soon; mistral-small3.2; and all-minilm and other embedding models.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
Ollama now includes a significantly improved model scheduling system, reducing crashes due to out of memory issues, maximizing GPU utilization and performance, especially on multi-GPU systems.