Data Machina #247
Data Machina
Open source mixture-of-experts models from AI21Labs, Alibaba, MetaAI, Databricks, and xAI are achieving near state-of-the-art performance comparable to closed models from OpenAI and Google. Databricks' DBRX uses 132B total parameters with 36B active per input, Alibaba's Qwen1.5-MoE-A2.7B reduces training costs by 75% while matching 7B model performance, and xAI's Grok-1.5 features a 128K context window. These efficient open MoE architectures allow researchers and developers to deploy competitive alternatives to proprietary large language models with reduced computational overhead.
Why it matters
New Open Mixture-of-Experts Models. Jamba SSM-MoE. Qwen1.5-MoE-A2.7B. DBRX 132B MoE. frankenMoEs. AI Agentic Workflows. 1-bit ML Models. OpenDevin. AgentStudio.