TLDRocket
Sign in

EMO: Pretraining mixture of experts for emergent modularity

Allen Institute (AI2)

Researchers released EMO, a mixture-of-experts language model with 128 total experts and 1 billion active parameters trained on 1 trillion tokens, designed so that modular structure emerges naturally from data without predefined domains. The model can use just 12.5% of its experts for specific tasks while retaining 97% of full-model performance, whereas standard MoE models degrade sharply when experts are pruned. EMO achieves this by restricting tokens within the same document to activate from a shared expert pool during training, causing experts to organize around semantic domains like health and politics rather than surface features like prepositions.

Why it matters

EMO is a new mixture-of-experts model trained so modular expert groups emerge from data, enabling users to select small task-specific expert subsets while preserving near full-model performance.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.