Ollama now supports Jev-style decision models
Ollama 0.35 adds support for Jev-style decision models via a new /v1/systemone endpoint, with three models available today: Nimble 9B from Bespoke Labs and two experimental Tev1 models from Together AI.
First-party releases and research from Ollama. Headlines and excerpts link to the original articles.
Ollama 0.35 adds support for Jev-style decision models via a new /v1/systemone endpoint, with three models available today: Nimble 9B from Bespoke Labs and two experimental Tev1 models from Together AI.
Ollama announced that its Pro, Max, and Team plans now use per-token pricing with monthly included usage credits, and introduced its Team plan at $500/month introductory pricing. Existing subscribers keep their current plans and can upgrade in account settings.
Ollama announced that Claude Desktop can now be configured to work with Ollama as a third-party gateway provider, allowing developers to run open models inside Claude. The integration supports both local and cloud models, with telemetry disabled by default and a Zero Data Retention policy.
NVIDIA Nemotron 3.5 Lightning, a 30 billion parameter open model with 3B active parameters, is now available on Ollama and runs entirely on local devices. It is designed for long-running agentic tasks such as tool calling and multi-step work.
Meta Superintelligence Labs has released Muse Glimmer, a 30B multimodal open model under the Apache 2.0 license, now available on Ollama for local coding agents and personal assistants. Ollama's MLX engine adds DFlash and image input support, running Muse Glimmer 1.5x-1.8x faster on Apple Silicon.
Ollama announced it has raised $88M from Benchmark, Theory Ventures, 8VC, Y Combinator and angel investors. The company says it serves 8.9 million developers and is used by 85% of the Fortune 500.
Ollama 0.31 makes Gemma 4 significantly faster on Apple Silicon using multi-token prediction powered by MLX, reporting up to 90% faster token generation on the Aider polyglot benchmark. The speedup is enabled by default and does not change the model's output.
Ollama updated its MLX engine for Apple Silicon, adding support for NVIDIA's NVFP4 quantization format and optimizations it says deliver up to 20% faster output, higher-quality responses, and lower memory use. The release also introduces a snapshot system for agent workloads.
Ollama 0.30 adds GGUF model compatibility through llama.cpp, up to 20% faster performance on NVIDIA hardware, and Vulkan enabled by default for broader GPU support on AMD and Intel devices.
NVIDIA Nemotron 3 Ultra, a 550 billion parameter open model with 55B active, is now available on Ollama's cloud for long-running agentic workflows. It offers 1M token context and is optimized for NVIDIA's NVFP4 format, with Ollama citing up to 30% cost savings versus other leading open models.
OpenJarvis v1.0, an open-source framework for building personal AI agents that run on local hardware, is now available with built-in support for Ollama. It is built by Stanford's Hazy Research and Scaling Intelligence labs as part of their Intelligence Per Watt research into efficient local AI.
Ollama has released a preview build powered by Apple's MLX framework on Apple Silicon, citing large speedups on all Apple Silicon devices and support for NVIDIA's NVFP4 format. The preview, Ollama 0.19, accelerates the Qwen3.5-35B-A3B model and requires a Mac with more than 32GB of unified memory.
Ollama 0.17 can install and configure OpenClaw, a personal AI assistant, with the single command ollama launch openclaw --model kimi-k2.5:cloud. Setup requires Ollama 0.17 or later, Node.js, and a Mac or Linux system, with Windows via WSL.
Ollama announced support for subagents and built-in web search in Claude Code, requiring no MCP servers or API keys. Subagents run tasks in parallel, and web search is integrated into Ollama's Anthropic compatibility layer, working with any model on Ollama's cloud.
OpenClaw is a personal AI assistant that links messaging platforms to local AI coding agents through a centralized gateway running on the user's own devices. Ollama published installation and launch instructions, including an ollama launch openclaw command and a recommended context length of at least 64k tokens.
Ollama released a new command, ollama launch, that sets up and runs coding tools such as Claude Code, OpenCode and Codex with local or cloud models without environment variables or config files. It requires Ollama v0.15+ and recommends at least 64000 tokens of context length.
Ollama has added experimental image generation support on macOS, with Windows and Linux planned. The feature runs models like Alibaba's Z-Image Turbo and Black Forest Labs' FLUX.2 Klein locally via the ollama run command.
Ollama v0.14.0 and later are now compatible with the Anthropic Messages API, allowing tools like Claude Code to run with open-source models locally or via ollama.com. The release supports tool calling, streaming, vision, and other features.
Open models can now be used with OpenAI's Codex CLI through Ollama, according to an Ollama post dated January 15, 2026. Codex can read, modify, and execute code in the working directory using models such as gpt-oss:20b, gpt-oss:120b, or other open-weight alternatives.
Ollama is partnering with OpenAI and ROOST to release the gpt-oss-safeguard reasoning models for safety classification, available in 20B and 120B sizes under the Apache 2.0 license.