Cloud models
Ollama announced on September 19, 2025 that cloud models are now in preview. According to the company, the preview lets users run larger models with fast, datacenter-grade hardware, while continuing to use their local tools. Ollama states that its cloud does not retain user data, citing privacy and security.
The company says the same Ollama experience now works across local and cloud environments and integrates with existing tools. Cloud models also work through Ollama's OpenAI-compatible API.
Four models are listed as available in the preview: qwen3-coder:480b-cloud, gpt-oss:120b-cloud, gpt-oss:20b-cloud, and deepseek-v3.1:671b-cloud.
Getting started requires downloading Ollama v0.12 and running a command in a terminal, such as ollama run qwen3-coder:480b-cloud. Ollama says cloud models behave like regular models, so users can ls, run, pull, and cp them as needed.
Usage examples are given for Ollama's JavaScript library, installed with npm install ollama, and its Python library, installed with pip install ollama. In both cases, users pull a cloud model first, such as gpt-oss:120b-cloud, then call chat. A cURL example calls the endpoint at localhost:11434/api/chat with the same model.
Ollama says cloud models use inference compute on ollama.com and require signing in to ollama.com via ollama signin; users can run ollama signout to stay signed out. Cloud models can also be accessed directly through ollama.com's API.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
Cloud models are now in preview, letting you run larger models with fast, datacenter-grade hardware. You can keep using your local tools while running larger models that wouldn’t fit on a personal computer.