TLDRocket
Sign in
Latest Google releases a new local-first Granola competitor — TechCrunch China’s Manus raises over $500M in first funding round since split wit... — TechCrunch Musk’s Grok Bot Turns to Claude for Help — Trending Topics Ethereum Researcher Warns A.I. Could Break Blockchain Encryption Befor... — Trending Topics Google Cloud introduces Gemini agent to change enterprise work — SiliconANGLE GPT-6 for Everyone: ChatGPT Now Answers With Buttons, Maps and Mini-Ap... — Trending Topics Nvidia's big bet on physical AI aims for safer robotaxis, humanoid rob... — Ars Technica A senator tried to ban gambling on prediction markets—now she's a Kals... — Ars Technica

The AI intelligence platform

Every AI story that matters — and the intelligence behind it.

TLDRocket reads all relevant sources, removes duplicate coverage, and publishes a short neutral summary of every story, linking back to the original. Free, no spam, unsubscribe anytime.

Add to Slack

Every story also updates live profiles event timelines weekly rankings the AI Market Index

Thursday, 28 March 2024

Qwen1.5-MoE: Matching 7B Model Performance with 1/3 Activated Parameters

GitHub Pages 2 years ago 53

Alibaba's Qwen team released Qwen1.5-MoE-A2.7B, a mixture-of-experts language model with 2.7 billion activated parameters that matches the performance of 7-billion-parameter models like Mistral 7B and Qwen 1.5-7B on benchmarks including MMLU, GSM8K, and HumanEval. The model reduces training costs by 75% and achieves 1.74 times faster inference speed compared to the standard Qwen 1.5-7B, while using only one-third of the non-embedding parameters. The efficiency gains enable practitioners to deploy comparable model quality with substantially lower computational requirements for both training and inference.

Mamba Explained

The Gradient 2 years ago 32

Mamba is a State Space Model architecture that replaces the Attention mechanism in Transformers with a control-theory-inspired SSM while maintaining similar performance and scaling properties. A Mamba-3B model outperforms same-sized Transformers and matches Transformers twice its size, while running up to 5 times faster and handling sequence lengths up to 1 million tokens. This enables faster inference and longer context windows by eliminating the quadratic time complexity bottleneck that limits Transformer efficiency.

Making education data accessible

OpenAI 2 years ago 18

Zelma is using GPT-4 to help make education data more accessible to users. The system processes educational datasets and extracts information through natural language interactions rather than requiring manual database queries. This approach reduces the technical barriers for educators and administrators to access the information they need from education records.

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.