MiniMax M2
MiniMax M2 is now available on Ollama's cloud, described as a model built for coding and agentic workflows. Ollama also published setup instructions for using it with tools including VS Code, Zed, and Droid, plus cloud API access.
First-party releases and research from Ollama. Headlines and excerpts link to the original articles.
MiniMax M2 is now available on Ollama's cloud, described as a model built for coding and agentic workflows. Ollama also published setup instructions for using it with tools including VS Code, Zed, and Droid, plus cloud API access.
Ollama published benchmark results for running various models on NVIDIA DGX Spark hardware using firmware 580.95.05 and Ollama v0.12.6, reporting prefill and decode tokens-per-second figures across models including gpt-oss, gemma3, llama3.1, deepseek-r1, and qwen3.
Ollama announced on October 16, 2025 that GLM-4.6 and Qwen3-coder-480B are available on its cloud service with integrations for coding tools, and that Qwen3-Coder-30B has been updated for faster, more reliable tool calling in its new engine.
Ollama announced a partnership with NVIDIA for the NVIDIA DGX Spark, saying Ollama runs fast and efficiently out-of-the-box on the device. The DGX Spark is powered by the NVIDIA GB10 Grace Blackwell Superchip and delivers 1 petaFLOP of performance with 128GB of memory.
Ollama released a web search API on September 24, 2025, alongside a web fetch API. A free tier is available for individuals, with higher rate limits via Ollama's cloud, plus REST support and Python and JavaScript library integrations.
Ollama released a new model scheduling system that measures exact memory requirements before running a model, aiming to cut out-of-memory crashes and improve GPU utilization, including on multi-GPU and mismatched GPU systems. It is enabled by default for models on Ollama's new engine.
Ollama announced on September 19, 2025 that cloud models are now in preview, allowing users to run larger models on datacenter-grade hardware. The cloud offering integrates with existing local tools and Ollama's OpenAI-compatible API, and Ollama says its cloud does not retain user data.
Ollama has partnered with OpenAI to bring OpenAI's gpt-oss open weight models, in 20B and 120B sizes, to Ollama and its community. Ollama supports the models natively in the MXFP4 quantization format and is collaborating with NVIDIA for acceleration on RTX GPUs.
Ollama released a new app for macOS and Windows on July 30, 2025, adding a way to download and chat with models, drag-and-drop file support, image input for compatible models, and code file processing. Standalone CLI downloads remain available on Ollama's GitHub releases page.
Ollama announced Secure Minions, a protocol built by Stanford's Hazy Research lab that encrypts communication between local Ollama models and frontier cloud models using NVIDIA H100 confidential computing, with under 1% added latency on long prompts.
Ollama added the ability to enable or disable model thinking, separating reasoning from output when enabled. DeepSeek R1 and Qwen 3 support the feature, available through CLI flags, interactive commands, a scripting option, and a new think parameter in the generate and chat APIs.