TLDRocket
Sign in

Ai2 releases Olmo-core 3 to make developing large mixture-of-experts LLMs more efficient

SiliconANGLE Kyt Dotson ● Covered by 3 sources

Ai2 says its new Olmo-core 3 makes big mixture-of-experts models train faster and cheaper. It can scale to over one trillion parameters without lighting up every GPU.

Based on reporting by SiliconANGLE, Kyt Dotson — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

The Allen Institute for AI says it has built a training framework that could make giant mixture-of-experts models a lot less painful to work on. The system, called Olmo-core 3, is meant to push MoE training to the trillion-parameter scale while keeping the compute bill under control.

That matters because MoE models are strange beasts. Instead of firing up the whole model for every token, they route work through a small set of specialist experts. Ai2 says its new setup is designed to bridge the gap between dense models and MoE systems, with the expert pool growing from eight to 128 while still picking only four experts per token.

The company’s benchmark numbers are the sharpest part of the pitch. On Nvidia B3000 GPUs, Olmo-core 3 processed 52,000 tokens per second on a 47-billion-parameter model. Ai2 compared that with Nvidia’s Megatron-core training architecture, which it said topped out around 19,400 tokens per second. By that measure, Olmo-core 3 delivered roughly 2.7 times the throughput.

Ai2 says the speedup comes from how the framework spreads the work around. Expert parallelism keeps only part of the expert pool on each GPU. Layer splitting reduces how much of the model each GPU has to hold. A distributed optimizer spreads optimizer state across multiple GPUs instead of duplicating it everywhere. The result is lower memory overhead as models grow, which is the whole game when the model itself and the training state are both getting huge.

The framework also supports MXFP8, a format that uses fewer bits for some values and can cut both compute and GPU traffic. Ai2 says the point is to give researchers a way to build and train larger models without needing state- or enterprise-level infrastructure. The project is already available on GitHub for developers and the open-source community.

My take — AI-written commentary, not fact-checked reporting

This is the kind of AI work that actually matters: less theater, more plumbing. Bigger models are easy to brag about; making them train without setting fire to the GPU budget is the unglamorous part everybody pretends not to notice. Open tooling here is the right move, because the field has enough secret sauce and not nearly enough useful roads.

Read more about this at: SiliconANGLE

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.