Introducing Olmo-core 3: Open, Scalable Training Infrastructure for Large MoEs
Published October 1, 2026
Overview
The Allen Institute for AI (AI2) has released Olmo-core 3, the latest iteration of its open training infrastructure, now purpose-built to handle the unique demands of large-scale Mixture-of-Experts (MoE) models. As MoE architectures become the dominant paradigm for frontier open-weight models in 2026, Olmo-core 3 aims to lower the barrier for researchers and organizations seeking to train their own high-performance models without relying on closed, proprietary stacks.
Why MoEs?
Mixture-of-Experts models activate only a subset of their parameters for each token, enabling dramatically larger total parameter counts at a fraction of the per-token compute cost of dense models. This efficiency has made MoEs the architecture of choice for state-of-the-art systems, but it also introduces significant engineering challenges:
- Load balancing across experts during training
- Communication overhead from all-to-all routing between devices
- Memory pressure from storing large expert parameter sets
- Checkpointing complexity for models with heterogeneous parameter distributions
Olmo-core 3 addresses each of these challenges directly.
What's New in Olmo-core 3
Scalable Expert Parallelism
Olmo-core 3 introduces improved expert parallelism strategies, allowing models to scale across thousands of GPUs while minimizing communication bottlenecks. The framework dynamically balances token routing to prevent expert collapse and underutilization.
Modular Training Pipeline
The training stack has been reorganized into modular components, making it easier to swap in custom data loaders, optimizers, or parallelism strategies without rewriting the entire training loop.
Robust Checkpointing
New checkpointing utilities support efficient, fault-tolerant saving and resumption of MoE training runs—critical for long training jobs that may span weeks or months.
Fully Open
Consistent with AI2's mission, Olmo-core 3 is released under a permissive open-source license, with reproducible recipes, documentation, and reference configurations available to the community.
Getting Started
Developers can access Olmo-core 3 through AI2's public repositories. The release includes example training configurations for MoE models at multiple scales, along with guidance on adapting the infrastructure to custom hardware setups.
Looking Ahead
As the open-source community continues to push the boundaries of what's possible with MoEs, tools like Olmo-core 3 play a critical role in ensuring that cutting-edge training infrastructure remains accessible to everyone—not just well-resourced labs. AI2 invites feedback and contributions as the framework evolves.
For more details, code, and documentation, visit the official AI2 blog and repository.
