AivexaNewsSearch
AI news for builders and product teamsChecked every hour

NVIDIA Nemotron 3 Ultra

Collected Oct 1, 2026

NVIDIA Nemotron 3 Ultra is now available on Ollama's cloud, according to Ollama. The model is a 550 billion parameter open model from NVIDIA with 55 billion active parameters, described as built for long-running, agentic workflows with fast and affordable performance across hundreds of tool calls.

Ollama lists several model highlights. It says the model is tuned for agent orchestration, coding agents, deep research, and complex enterprise workflows that run across hundreds of steps. It also states a 1M token context, intended to keep entire codebases, long tool histories, and research trails in context.

On efficiency, Ollama says the model uses 550B total parameters with only 55B active per token, and is optimized for NVFP4, NVIDIA's 4-bit floating point format, which packs the model into less memory and runs faster.

Ollama provides launch commands for several integrations: Claude Code via "ollama launch claude --model nemotron-3-ultra:cloud", Hermes Agent via "ollama launch hermes --model nemotron-3-ultra:cloud", OpenClaw via "ollama launch openclaw --model nemotron-3-ultra:cloud", and general chat via "ollama run nemotron-3-ultra:cloud". It directs users to download Ollama and says more integrations are available.

According to Ollama, benchmark figures show Nemotron 3 Ultra leads on accuracy across agent productivity, instruction following, and long-context tasks while delivering leading throughput, saving up to 30% on costs compared to other leading open models. Ollama's figures describe the model as leading among open models on agentic benchmarks for agent productivity, coding, and instruction following, positioned with leading accuracy and leading throughput, and on the cost efficiency frontier. Ollama references an NVIDIA Nemotron 3 Ultra blog and its own model page. The announcement is dated June 4, 2026.

Read at Ollama

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

NVIDIA Nemotron 3 Ultra is built for high-throughput reasoning and long-running agent workflows.