AivexaNewsSearch
AI news for builders and product teamsUpdated Oct 10, 2026, 20:01 UTC

Ollama

First-party releases and research from Ollama. Headlines and excerpts link to the original articles.

Latest stories

Newest first

MiniMax M2

MiniMax M2 is now available on Ollama's cloud, described as a model built for coding and agentic workflows. Ollama also published setup instructions for using it with tools including VS Code, Zed, and Droid, plus cloud API access.

Read original
OllamaFirst partyResearch

NVIDIA DGX Spark performance

Ollama published benchmark results for running various models on NVIDIA DGX Spark hardware using firmware 580.95.05 and Ollama v0.12.6, reporting prefill and decode tokens-per-second figures across models including gpt-oss, gemma3, llama3.1, deepseek-r1, and qwen3.

Read original
OllamaFirst partyIndustry

Qwen3-VL

Ollama announced on October 14, 2025 that Alibaba's Qwen3-VL, described as the most powerful vision language model in the Qwen series, is now available on Ollama's cloud, with local availability to follow soon.

Read original
OllamaFirst partyIndustry

NVIDIA DGX Spark

Ollama announced a partnership with NVIDIA for the NVIDIA DGX Spark, saying Ollama runs fast and efficiently out-of-the-box on the device. The DGX Spark is powered by the NVIDIA GB10 Grace Blackwell Superchip and delivers 1 petaFLOP of performance with 128GB of memory.

Read original

Web search

Ollama released a web search API on September 24, 2025, alongside a web fetch API. A free tier is available for individuals, with higher rate limits via Ollama's cloud, plus REST support and Python and JavaScript library integrations.

Read original

New model scheduling

Ollama released a new model scheduling system that measures exact memory requirements before running a model, aiming to cut out-of-memory crashes and improve GPU utilization, including on multi-GPU and mismatched GPU systems. It is enabled by default for models on Ollama's new engine.

Read original

Cloud models

Ollama announced on September 19, 2025 that cloud models are now in preview, allowing users to run larger models on datacenter-grade hardware. The cloud offering integrates with existing local tools and Ollama's OpenAI-compatible API, and Ollama says its cloud does not retain user data.

Read original

OpenAI gpt-oss

Ollama has partnered with OpenAI to bring OpenAI's gpt-oss open weight models, in 20B and 120B sizes, to Ollama and its community. Ollama supports the models natively in the MXFP4 quantization format and is collaborating with NVIDIA for acceleration on RTX GPUs.

OllamaFirst partyProducts

Ollama's new app

Ollama released a new app for macOS and Windows on July 30, 2025, adding a way to download and chat with models, drag-and-drop file support, image input for compatible models, and code file processing. Standalone CLI downloads remain available on Ollama's GitHub releases page.

Read original

Thinking

Ollama added the ability to enable or disable model thinking, separating reasoning from output when enabled. DeepSeek R1 and Qwen 3 support the feature, available through CLI flags, interactive commands, a scripting option, and a new think parameter in the generate and chat APIs.

Read original