One Model Family, Two Gold-Level Results: Fine-Tuning Nemotron for IOI and IMO

Starting from Nemotron 3, teams used supervised fine-tuning (SFT), reinforcement learning (RL), and feedback-driven inference to build specialist systems that reached gold-medal level at both IMO 2026 and IOI 2026, according to results described in a Hugging Face blog post.
The IOI result came from a live, prospective run under the same time, internet-access, and submission constraints as human contestants. It was described as an unofficial, unsupervised benchmark and was not included in the official IOI ranking. The IMO system's submitted proofs were graded by official IMO graders.
For competitive programming, the teams curated 22,000 problems and generated synthetic reasoning traces to train two specialists. Nemotron-3-Nano-CC, with 30 billion total parameters and 3 billion active parameters, received both SFT and RL. Nemotron-3-Ultra-CC, with 550 billion total parameters and 55 billion active parameters, received SFT.
On IOI 2025, Nano improved from 130 points before post-training to 280 after SFT and 291 after RL. With GenCorrect, an iterative generate-evaluate-refine strategy, it reached 468 points, crossing the gold threshold of 438.3. Ultra-CC reached 502 points with the same test-time strategy. The competition-specific Ultra-CC system used for IOI 2026 scored 535.4 out of 600.
The IMO project started from Nemotron 3 Ultra, training one specialist with SFT and another with RL. The SFT corpus contained 414,890 quality-filtered examples across 15,818 unique proof problems, covering proof generation, refinement, verification, and meta-verification. The RL model was trained on 9,597 proof problems selected near the model's capability frontier. For each IMO problem, models generated candidate proofs, scored them, produced critiques, and refined the most promising attempts, with a separate high-compute stage selecting the final submission. The system worked in natural language, with no formal prover, external tools, or internet access, and scored 30 out of 42 points, including full credit on four of six problems, exceeding the official gold-medal threshold.
The post states the IMO 2026 collection includes the SFT and RL checkpoints, both training datasets, and Nemotron-IMO-Bench, a benchmark of 200 olympiad-level problems. The Nemotron-3-Ultra-CC model is available on Hugging Face.
Based on reporting from the original publisher. Visit the source for full context and later updates.