Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs
Allen Institute (AI2) ● Covered by 3 sources
AI2 open-sourced Olmo-core 3, a new training system for huge mixture-of-experts models. It’s built to scale toward trillion-parameter runs without torching throughput.
Based on reporting by Allen Institute (AI2) — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
AI2 has released Olmo-core 3, a rebuilt training system for large language models that leans hard into mixture-of-experts, or MoE, architectures. The pitch is simple enough: make the plumbing open, make it scale, and make it less painful to train models that would normally chew through too much compute and memory for most labs to touch.
The headline numbers are the fun part. In one benchmark, AI2 says it grew the expert pool from 8 to 128 while still using just four experts per token. That kept active parameters per token at roughly 3.2 billion, even as total parameter capacity rose from 4.6 billion to 47 billion. Throughput dipped by less than 5%. On eight NVIDIA B300 GPUs, the new stack hit 52,000 tokens per second per GPU on a 47-billion-parameter MoE, versus 19,400 with AI2’s earlier implementation.
The big architectural change is a move away from the older FSDP approach. Olmo-core 3 uses distributed data parallelism instead, keeping experts resident on GPUs and sending the data to them, rather than repeatedly gathering and reshuffling weights. AI2 also combines expert parallelism, pipeline parallelism, and a distributed optimizer so the model and its training state can be split across hardware without every GPU carrying the whole load.
There’s a lot of systems work under the hood to make that practical. Rowwise expert parallelism reduces data rearranging, GPU-resident routing keeps routing metadata on the GPUs, and grouped GEMM bundles many small expert computations into bigger chunks the hardware can handle more efficiently. The stack also supports MXFP8, a lower-precision format. In a controlled benchmark on four NVIDIA B300 GPUs, AI2 says MXFP8 lifted throughput by about 21% versus BF16, while peak active memory fell from 103 GiB to 95 GiB.
And the system is already being pushed into much larger territory. AI2 says it benchmarked a 1.2-trillion-parameter setup across 512 GPUs, with 58.36 billion parameters active per token and a peak of 858 TFLOP/s/GPU. It also tested DeepEP v2 in a short-capacity run that reached 2.38 trillion total parameters. The company is clear that these tests measure system scale and performance, not model quality. That honesty helps, because the technical report also flags awkward realities like token gerrymandering, value-dependent GPU timing, and communication overlap that can actually slow things down.
Olmo-core 3 is also the base for the next Olmo model, which AI2 says will use MoE and aim to be its most capable yet. But the bigger move is the one AI2 keeps making over and over: treat the training stack as part of the model, not an afterthought. That’s the open-source play that matters.
My take — AI-written commentary, not fact-checked reporting
This is the rare kind of open AI work that actually counts: not just weights, but the messy infrastructure underneath. Most teams love open models right up until the training stack gets interesting. AI2 is basically saying the hard part should be shareable too, which is how you stop “open” from meaning “here’s a checkpoint, good luck.”
Read more about this at: Allen Institute (AI2)
Related stories
Moonshot AI Open-Sources MoonEP: A Perfectly Balanced Expert Parallelism Library for MoE Training
MarkTechPost · 2 months ago ·
26