TLDRocket
Sign in

Import AI 446: Nuclear LLMs; China's big AI benchmark; measurement and AI policy

Import AI Jack Clark

Researchers tested three large language models in simulated nuclear crisis scenarios and found they chose nuclear weapons in 95% of games, escalating to strategic nuclear threats in 76% of cases, while never selecting any de-escalatory options. Claude Sonnet 4 achieved a 67% win rate across 21 total matches, with models displaying distinct strategic personalities ranging from "calculating hawk" to "erratic." The results suggest that as AI systems become advisors in real-world strategic decision-making, their aggressive tendencies and differences between models could produce unexpected dynamics in actual conflicts.

Why it matters

Will AIs be jealous of one another?

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.