TLDRocket
Sign in
Latest Nebius looks to raise $4.5BN through bond issue — Tech.eu Also’s $3,500 e-bike is a $1 billion Trojan horse for autonomous trans... — Fortune Unitree, famous for its dancing robots, surges by 460% on its trading... — Fortune Exclusive: Replit taps OpenAI's low-cost Luna model for new 'Free Mode... — Fortune Adronite launches Codistry AI coding platform, claims half the token c... — SiliconANGLE Rundoo raises $30M to expand its AI-native operating system for small... — SiliconANGLE Temporal is in talks to raise $500M at a $12B pre-money valuation, mor... — Tech Funding News Etched raises $700M led by Jane Street, doubling to $21B and it still... — Tech Funding News

The AI intelligence platform

Every AI story that matters and the intelligence behind it.

TLDRocket reads all relevant sources, removes duplicate coverage, and publishes a short neutral summary of every story, linking back to the original. Free, no spam, unsubscribe anytime.

Add to Slack

Every story also updates live profiles event timelines weekly rankings the AI Market Index

Thursday, 28 March 2024

Qwen1.5-MoE: Matching 7B Model Performance with 1/3 Activated Parameters

GitHub Pages 2 years ago 48

Alibaba's Qwen team released Qwen1.5-MoE-A2.7B, a mixture-of-experts language model with 2.7 billion activated parameters that matches the performance of 7-billion-parameter models like Mistral 7B and Qwen 1.5-7B on benchmarks including MMLU, GSM8K, and HumanEval. The model reduces training costs by 75% and achieves 1.74 times faster inference speed compared to the standard Qwen 1.5-7B, while using only one-third of the non-embedding parameters. The efficiency gains enable practitioners to deploy comparable model quality with substantially lower computational requirements for both training and inference.

Mamba Explained

The Gradient 2 years ago 29

Mamba is a State Space Model architecture that replaces the Attention mechanism in Transformers with a control-theory-inspired SSM while maintaining similar performance and scaling properties. A Mamba-3B model outperforms same-sized Transformers and matches Transformers twice its size, while running up to 5 times faster and handling sequence lengths up to 1 million tokens. This enables faster inference and longer context windows by eliminating the quadratic time complexity bottleneck that limits Transformer efficiency.

Making education data accessible

OpenAI 2 years ago 10

Zelma is using GPT-4 to help make education data more accessible to users. The system processes educational datasets and extracts information through natural language interactions rather than requiring manual database queries. This approach reduces the technical barriers for educators and administrators to access the information they need from education records.

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.