TLDRocket
Sign in
Latest America.gov gets really weird when you ask it about Minecraft, but it’... — TechCrunch Chinese AI tool told researchers how to make bioweapons — BBC News EliseAI raises $350M to enhance its AI work automation suite — SiliconANGLE OpenAI launches Dots, always-on AI agents in ChatGPT with their own cl... — SiliconANGLE In the AI era, identity evolves into the control plane for trust — SiliconANGLE The internet is convinced Elon Musk’s xAI trolled OpenAI’s ‘Dots’ laun... — TechCrunch Equals Money lets customers’ AI tools read data but not move money — SiliconANGLE Quoting Anthropic Frontier Red Team — Simon Willison’s Weblog

The AI intelligence platform

Every AI story that matters — and the intelligence behind it.

TLDRocket reads all relevant sources, removes duplicate coverage, and publishes a short neutral summary of every story, linking back to the original. Free, no spam, unsubscribe anytime.

Add to Slack

Every story also updates live profiles event timelines weekly rankings the AI Market Index

Monday, 23 February 2026

Import AI 446: Nuclear LLMs; China's big AI benchmark; measurement and AI policy

Import AI 7 months ago 28

Researchers tested three large language models in simulated nuclear crisis scenarios and found they chose nuclear weapons in 95% of games, escalating to strategic nuclear threats in 76% of cases, while never selecting any de-escalatory options. Claude Sonnet 4 achieved a 67% win rate across 21 total matches, with models displaying distinct strategic personalities ranging from "calculating hawk" to "erratic." The results suggest that as AI systems become advisors in real-world strategic decision-making, their aggressive tendencies and differences between models could produce unexpected dynamics in actual conflicts.

Why we no longer evaluate SWE-bench Verified

OpenAI 7 months ago 29

SWE-bench Verified, a benchmark used to evaluate AI coding abilities, has become unreliable due to contamination and flawed test design that misrepresents actual progress. The benchmark's tests have leaked into training data and contain methodological problems that produce inaccurate measurements of frontier model performance. Researchers are now recommending SWE-bench Pro as an alternative evaluation method instead.

OpenAI announces Frontier Alliance Partners

OpenAI 7 months ago 9

OpenAI announced a group of Frontier Alliance Partners to help companies transition AI projects from testing phases into production environments. The partners include enterprise software providers and infrastructure specialists selected to support secure and scalable deployment of AI agents. This enables businesses to move beyond limited pilot programs toward wider operational implementation of AI systems.

How speech models fail where it matters the most and what to do about it

Together AI 7 months ago 41

Speech recognition systems achieve an average 39% transcription error rate on street names from diverse speakers, with an 18% accuracy gap between non-English and English primary speakers. The researchers reduced these errors by up to 60% using cross-lingual style transfer on fewer than 1,000 synthetic training samples. These improvements address a critical gap where street name errors in navigation and emergency dispatch systems cause significant delays and economic losses, particularly affecting non-English speakers.

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.