TLDRocket
Sign in
Latest Cambridge startup Neela Biotech raises £2.1M from Elbow Beach to turn... — Tech Funding News Ireland’s top-funded tech companies in H1 2026 — Tech.eu Researchers Suspect OpenAI and Anthropic of Stealing Their Discoveries — Trending Topics Search Engine Ecosia Swaps Mistral for Chinese A.I. — Trending Topics UK and Germany launch industrial tech corridor to accelerate AI and de... — Tech.eu University of Sussex spinout Universal Quantum raises “record” $100M S... — Tech.eu The Next Top Open Model, Google Voice Agents, DeepSeek Shrinks Caches:... — The Batch Dutch startup Innercrowd raises €750K to turn festival fans into marke... — Tech.eu

The AI intelligence platform

Every AI story that matters — and the intelligence behind it.

TLDRocket reads all relevant sources, removes duplicate coverage, and publishes a short neutral summary of every story, linking back to the original. Free, no spam, unsubscribe anytime.

Add to Slack

Every story also updates live profiles event timelines weekly rankings the AI Market Index

Sunday, 31 March 2024

Data Machina #247

Substack 2 years ago 33

Open source mixture-of-experts models from AI21Labs, Alibaba, MetaAI, Databricks, and xAI are achieving near state-of-the-art performance comparable to closed models from OpenAI and Google. Databricks' DBRX uses 132B total parameters with 36B active per input, Alibaba's Qwen1.5-MoE-A2.7B reduces training costs by 75% while matching 7B model performance, and xAI's Grok-1.5 features a 128K context window. These efficient open MoE architectures allow researchers and developers to deploy competitive alternatives to proprietary large language models with reduced computational overhead.

Task-Specific LLM Evals that Do & Don't Work

Eugene Yan 2 years ago 26

The article discusses evaluation metrics and methods for assessing large language model performance on specific tasks including classification, extraction, summarization, and translation. Key concrete metrics mentioned are ROC-AUC and PR-AUC for classification (ranging from 0.0 to 1.0), natural language inference models for measuring factual consistency in summaries, and specialized tools like chrF and COMET for translation quality. The author recommends moving beyond generic off-the-shelf evaluations toward task-specific metrics that better correlate with actual application performance and can reliably measure production-ready systems.

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.