Olmo-core 3 scales mixture-of-experts training to trillion parameters
Hugging Face releases Olmo-core 3, an open infrastructure that boosts MoE training throughput and supports models with over one trillion total parameters.
Hugging Face has released Olmo-core 3, a major update to its open-source framework for training large language models. Published on October 1, 2026, this version introduces a redesigned system specifically optimized for mixture-of-experts (MoE) architectures, enabling efficient scaling into the trillion-parameter range.
What happened
The release addresses a critical bottleneck in modern AI development: the high cost and energy consumption of training massive models. While MoE models offer theoretical efficiency by activating only a subset of parameters for each input, the overhead of coordinating these experts across GPU clusters often erodes those gains. Olmo-core 3 aims to close this gap by rethinking how data and model weights are distributed during training.
In internal benchmarks, the new framework demonstrated significant performance improvements. When increasing the expert pool from 8 to 128 while keeping active parameters per token fixed at approximately 3.2 billion, total parameter capacity grew from 4.6 billion to 47 billion. Despite this massive increase in model size, training throughput dropped by less than 5 percent. The infrastructure has also been tested at scales exceeding one trillion total parameters.
This evolution marks a shift from previous iterations. While Olmo 3 used a dense architecture where nearly all parameters were active for every token, Olmo-core 3 is built from the ground up for sparse MoE designs. It replaces the earlier fully sharded data parallelism approach with a system based on distributed data parallelism, keeping experts resident on GPUs to avoid repeated weight gathering.
How it works
Olmo-core 3 distributes the model and its training state across hardware using three primary techniques. Expert parallelism spreads the specialized components across different GPUs, so each device stores only a fraction of the total expert pool. Pipeline parallelism splits the model layers across GPU groups, reducing memory requirements per device. A distributed optimizer further saves memory by spreading the optimizer state across the cluster instead of replicating it on every GPU.
To minimize communication overhead, the framework employs several optimizations. Rowwise expert parallelism places routed data directly into expert input buffers, reducing data rearrangement costs. GPU-resident routing keeps metadata on the graphics processors, allowing the CPU to queue work without waiting for data transfers. Additionally, grouped GEMM operations combine many small computations into larger batches, improving GPU execution efficiency.
The system also supports MXFP8, a lower-precision number format that reduces data movement and computation costs. In benchmarks on NVIDIA B300 GPUs, enabling MXFP8 increased training throughput by about 21 percent compared to the BF16 baseline, while also reducing peak active memory usage. These features work together to balance trade-offs between computation speed and data transfer latency.
Key details
- Olmo-core 3 achieved 2.7 times higher throughput than its predecessor in preliminary tests on eight NVIDIA B300 GPUs.
- The framework supports models with over one trillion total parameters, demonstrated in a test with 58.36 billion active parameters per token.
- Using MXFP8 precision improved throughput by 21 percent and reduced peak memory from 103 GiB to 95 GiB in controlled benchmarks.
- The system identifies a phenomenon called "token gerrymandering," where routing scores improve even as workload balance worsens.
- Overlapping communication and computation on separate GPU streams did not always increase speed and sometimes slowed execution.
- Lowering learning rates for experts did not improve results in the tested model families, contrary to some common assumptions.
Why it matters
For engineering teams building large-scale AI systems, Olmo-core 3 provides a transparent alternative to proprietary training stacks like NVIDIA’s Megatron-Core. By open-sourcing the infrastructure behind its next-generation models, Hugging Face allows researchers and smaller labs to experiment with MoE architectures without being locked into specific vendor ecosystems. This transparency is crucial for understanding the real-world performance characteristics of sparse models.
The technical findings included in the release highlight important pitfalls for developers. For instance, the discovery that overlapping communication and computation can sometimes degrade performance challenges standard optimization heuristics. Similarly, the observation that input values affect computation time even when matrix dimensions remain constant suggests that benchmarking must be done with care. These insights help engineers avoid costly misconfigurations when scaling their own models.
What you can do
- Review the Olmo-core 3 technical report to understand the specific ablations and design choices made during development.
- Experiment with the open-source code on GitHub to test MoE training configurations on your own hardware.
- Use the interactive walkthrough provided by Hugging Face to visualize how data, expert, and pipeline parallelism interact.
- Benchmark your current training stack against the reported metrics, ensuring you match input values for fair comparisons.
- Investigate the impact of MXFP8 precision on your specific workload to determine if the conversion overhead is worth the throughput gain.
- Monitor for "token gerrymandering" in your routing metrics to ensure that load balancing scores reflect actual workload distribution.


