Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs
Hugging Face announced the release of Olmo-core 3, described as a significant upgrade to its framework for developing large language models, featuring a redesigned open mixture-of-experts (MoE) training system. The company said it is designed to scale MoE training into the trillion-parameter range while preserving computational efficiency, and that it is one of the core systems behind the next generation of Olmo.
In one benchmark, the expert pool was increased from 8 to 128 while still selecting only four experts per token, keeping active parameters per token roughly fixed at about 3.2B. Total parameter capacity grew from 4.6B to 47B, while training throughput fell by less than 5%. The infrastructure has also been benchmarked at over one trillion total parameters.
Olmo-core 3 switches from fully sharded data parallelism (FSDP) to a system based on distributed data parallelism (DDP), keeping experts resident on GPUs and routing data to them. In a preliminary test on eight NVIDIA B300 GPUs, a 47-billion-parameter MoE processed 52,000 tokens per second per GPU with the new stack, compared with 19,400 using the earlier implementation, about 2.7 times the throughput.
The framework combines expert parallelism, pipeline parallelism, and a distributed optimizer, along with rowwise expert parallelism, GPU-resident routing, and grouped GEMM. It also supports MXFP8. In a controlled benchmark on four NVIDIA B300 GPUs, enabling MXFP8 where it helped most raised training throughput about 21% over BF16, while peak active memory fell from 103 GiB to 95 GiB.
Hugging Face benchmarked a 1.2-trillion-parameter model with 58.36 billion parameters active per token across 512 GPUs, with highest observed throughput of 858 TFLOP/s/GPU. It also experimented with DeepEP v2, reaching a configuration with 2.38 trillion total parameters in a short-capacity test, which it said demonstrates scale rather than sustained training performance.
Based on reporting from the original publisher. Visit the source for full context and later updates.