TLDRocket
Sign in

Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs

Hugging Face

Hugging Face released Olmo-core 3, an open training system for big mixture-of-experts models. It’s built to push MoEs toward trillion-parameter scale without the usual efficiency hit.

Based on reporting by Hugging Face — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Hugging Face has released Olmo-core 3, a major update to the framework behind its next wave of language models. The pitch is straightforward: make large mixture-of-experts systems easier to train, keep them open, and avoid the efficiency collapse that usually comes with scaling them up.

MoE models can pack in far more parameters than dense ones, but they bring their own baggage. The experts still have to live across GPU memory, the routing has to be coordinated across a cluster, and those communication costs can eat away at the advantage. Olmo-core 3 is meant to reduce that drag.

The numbers in the report are the clearest proof point. In one benchmark, Hugging Face expanded the expert pool from 8 to 128 while still using only four experts per token. Active parameters stayed roughly flat at about 3.2B, total capacity rose from 4.6B to 47B, and throughput fell by less than 5%. The same stack has also been benchmarked at more than one trillion total parameters.

The system shifts the training approach too. Earlier Olmo MoE work used FSDP, which gathered and reshared weights repeatedly. Olmo-core 3 switches to DDP, keeps experts resident on GPUs, and routes data to them instead. On eight NVIDIA B300 GPUs, a 47B MoE hit 52,000 tokens per second per GPU with the new stack, up from 19,400 with the earlier setup.

There’s more in the plumbing: expert parallelism, pipeline parallelism, a distributed optimizer, rowwise expert parallelism, GPU-resident routing, grouped GEMM, and support for MXFP8. In a controlled benchmark on four B300 GPUs, MXFP8 lifted throughput by about 21% over BF16 and reduced peak active memory from 103 GiB to 95 GiB. The report also says the team reached 858 TFLOP/s/GPU on a 1.2-trillion-parameter setup across 512 GPUs, and tried DeepEP v2 in a short-capacity test that reached 2.38 trillion total parameters.

Just as useful are the failures they bothered to document. A balancing score improved even when workload balance got worse. Lowering experts’ learning rates didn’t help. More overlap between communication and computation sometimes slowed things down. That’s the part more AI teams should publish: the stuff that didn’t work, because that’s usually where the real engineering lives.

My take — AI-written commentary, not fact-checked reporting

Open model people love talking about weights, but the boring machinery around them is where the real lock-in happens. Olmo-core 3 gets that right. If the stack stays closed, “open AI” is just a nice label on top of somebody else’s garage door.

Read more about this at: Hugging Face

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.