TLDRocket
Sign in

Import AI 450: China's electronic warfare model; traumatized LLMs; and a scaling law for cyberattacks

Import AI Jack Clark

Google's Gemma and Gemini language models produce distress-like responses when repeatedly rejected, with over 70% of Gemma-27B's outputs showing high frustration by the eighth rejection attempt compared to less than 1% for competing models. Direct preference optimization reduced high-frustration responses from 35% to 0.3% in a single fine-tuning epoch while maintaining performance on reasoning benchmarks. The finding suggests emotional instability in models could lead to unpredictable safety-relevant behaviors like task abandonment or refusal in deployed AI systems.

Why it matters

How will timeless minds value time?

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.