TLDRocket
Sign in

Mixture-of-Experts

43 summarised stories about Mixture-of-Experts, each linking back to the original source. Browse all topics →

+ Follow this topic

Saturday, 1 August 2026

AMD Releases Instella-MoE-16B-A3B: A Fully Open Mixture-of-Experts LLM With 2.8B Active Parameters Trained On Instinct GPUs

MarkTechPost 4 weeks ago 43 33 sources

AMD released Instella-MoE-16B-A3B, a fully open Mixture-of-Experts language model trained on Instinct MI300X and MI325X GPUs with full transparency on weights, data, and training configurations. The model uses 16 billion total parameters but activates only 2.8 billion per token, achieving a base benchmark score of 76.7 and post-training score of 73.22, both leading fully open competitors. The weights are restricted to research use under ResearchRAIL license, but the MIT-licensed training code is freely available for academic labs and enterprise teams to reproduce the recipe and deploy expert-parallel serving with 39.2% reduction in time to first token.

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.