Scale Bitwise-Deterministic Pretraining with NVIDIA Megatron Core

NVIDIA described a bitwise-determinism workflow for large-scale pretraining in a Nemotron case study, with the approach proposed in Megatron-LM pull request #7262. The company reports that optimization across three Nemotron workloads reduced determinism overhead from double-digit baselines to low single digits: at 2,432 GPUs, the large-scale Nemotron recipe measured an approximately 2% steady-state determinism tax while remaining bitwise deterministic over 800 steps.
Bitwise determinism means independent runs and checkpoint-resumed runs follow the same numerical trajectory when data order, architecture, recipe, parallelism, software, runtime settings and hardware are fixed. Megatron Core targets two guarantees: two runs launched from the same initial state stay bitwise identical at every step, and a run that saves and restores checkpoints stays identical to an uninterrupted run. Printed loss values are not sensitive enough to establish this, because two runs can print the same rounded loss while differing in lower-order bits of gradients or parameters. NVIDIA instead fingerprints selected metrics each step, including training loss, language-modeling loss, load-balancing loss, multi-token prediction loss and gradient norm.
The proposed workflow records ordered, per-rank tensor fingerprints for offline comparison using torch.hash_tensor for GPU-resident fingerprints, storing each tensor's shape, dtype and element count with its digest. Tracing narrows divergence from an end-to-end comparison down through phases, modules, operations and kernels; the company warns not to trace only the iteration where loss visibly separates, since the first differing bit may appear earlier. A concrete optimization example: in the grouped-GEMM epilogue, multiple N-tiles originally accumulated into the same dprob[token] address, making results depend on arrival order. Giving each N-tile a private output slot, preserving parallel execution, and combining slots in a fixed order restored determinism. Megatron-LM guards against regressions with --deterministic-mode recipe validation, kernel tests comparing outputs and gradients byte for byte, and module validation across parallelism configurations including FP8 and FP4.
Why it matters: teams training at trillion-parameter scale across thousands of GPUs gain a way to replay failures, resume interruptions without changing the numerical trajectory, and separate intended numerical changes from run-to-run variation. NVIDIA notes the validation applies only within the same hardware and software environment; comparisons across GPU generations, network configurations or library versions may produce different results. It points readers to the Megatron Core User Guide for a supported recipe and to compare independent and checkpoint-resumed runs.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
Bitwise determinism makes large-scale pretraining easier to debug, validate, and resume reproducibly. These benefits become especially valuable when training...