TLDRocket
Sign in

Qwen

47 summarised stories about Qwen, each linking back to the original source. Browse all topics →

+ Follow this topic

Thursday, 28 March 2024

Qwen1.5-MoE: Matching 7B Model Performance with 1/3 Activated Parameters

GitHub Pages 2 years ago 48

Alibaba's Qwen team released Qwen1.5-MoE-A2.7B, a mixture-of-experts language model with 2.7 billion activated parameters that matches the performance of 7-billion-parameter models like Mistral 7B and Qwen 1.5-7B on benchmarks including MMLU, GSM8K, and HumanEval. The model reduces training costs by 75% and achieves 1.74 times faster inference speed compared to the standard Qwen 1.5-7B, while using only one-third of the non-embedding parameters. The efficiency gains enable practitioners to deploy comparable model quality with substantially lower computational requirements for both training and inference.

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.