NVIDIA Nemotron 3.5 Lightning
NVIDIA Nemotron 3.5 Lightning is now available on Ollama and runs completely on a user's own device, according to the Ollama post dated August 11, 2026. The model is a 30 billion parameter open model from NVIDIA with 3B active parameters, described as built for agents that stay running: gathering context, calling tools, and working through multi-step tasks.
Ollama said the model targets agentic work such as reading a file, calling a tool, sorting a result, and retrying a failed step, and that at 3B active parameters per token it is built for local systems rather than the datacenter. It uses a hybrid Mixture-of-Experts architecture. Ollama lists supported local hardware as NVIDIA RTX PCs, NVIDIA RTX PRO workstations, NVIDIA DGX Spark and DGX Station, plus the datacenter and cloud. The model has a context length up to 1M tokens, which Ollama says leaves room for long tool histories across multi-turn workflows.
Ollama said the model was developed with the Nemotron Coalition and trained for coding, tool calling, instruction following and multi-turn work. It supports speculative decoding using multi-token prediction (MTP), DFlash or DSpark, which Ollama says offers up to 4x higher throughput than comparable open models. Ollama states the model offers 4x higher throughput and 30% faster task completion time compared to other leading open models of similar size, with full results and test configurations in NVIDIA's launch blog.
Ollama lists example workloads including long-running personal assistants, coding sub-agents, security operations, a local tier alongside a cloud model, and post-training a specialist. The model is open weights and trained on open datasets. Ollama lists run commands for general chat, Claude Code, OpenClaw, Hermes Agent and OpenCode, and notes an MLX variant, nemotron-3.5-lightning:30b-mlx, for Apple silicon.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
NVIDIA Nemotron 3.5 Lightning is now available on Ollama. It's a 30 billion parameter (3B active) open model built for agents that stay running, gathering context, calling tools, and working through multi-step tasks on your own hardware.