NVIDIA DGX Spark performance
Ollama published performance test results for NVIDIA DGX Spark on October 23, 2025. The tests used release day firmware version 580.95.05 and Ollama v0.12.6 to measure how Ollama performs on the device.
Each test was run 10 times with temperature set to 0, output constrained to 500 tokens, and caching disabled so repeated tests would not be faster. The prompt asked for an in-depth summary of a story drawn from the first 200 lines of pg98.txt, the book "A Tale of Two Cities." Ollama said the test script and its readme are available and can be customized for individual testing.
Reported results for the DGX Spark include: gpt-oss 20B MXFP4 at 3.224k prefill and 58.27 decode tokens per second; gpt-oss 120B MXFP4 at 1.169k prefill and 41.14 decode; gemma3 12B q4_K_M at 1.894k prefill and 24.25 decode; gemma3 12B q8_0 at 1.406k prefill and 15.46 decode; gemma3 27B q4_K_M at 834.1 prefill and 10.83 decode; gemma3 27B q8_0 at 585.4 prefill and 7.210 decode; llama3.1 8B q4_K_M at 7.614k prefill and 38.02 decode; llama3.1 8B q8_0 at 6.110k prefill and 25.23 decode; llama3.1 70B q4_K_M at 1.911k prefill and 4.423 decode; deepseek-r1 14B q4_K_M at 5.919k prefill and 19.99 decode; deepseek-r1 14B q8_0 at 4.667k prefill and 13.32 decode; qwen3 32B q4_K_M at 705.0 prefill and 9.411 decode; and qwen3 32B q8_0 at 487.2 prefill and 6.240 decode.
Ollama noted that OpenAI's gpt-oss models are tested using models officially provided by OpenAI and distributed via Ollama. It said some GGUFs distributed online labeled as MXFP4 are further quantized to q8_0 in the attention layers, while the same layers are BF16 on Ollama as intended by OpenAI.
For firmware below 580.95.05, Ollama recommends updating through the DGX Dashboard, or via CLI by running apt update, apt dist-upgrade, fwupdmgr refresh, fwupdmgr upgrade, and reboot. Ollama also described installing Ollama via its install script and using OpenAI's Codex with the gpt-oss model, including the larger gpt-oss-120b model, which it said fits entirely into the 120GB of VRAM provided by the GB10 Grace Blackwell Superchip.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
We ran performance tests on release day firmware and an updated Ollama version to see how Ollama performs.