TLDRocket
Sign in

Adversarial Attacks

33 summarised stories about Adversarial Attacks, each linking back to the original source. Browse all topics →

+ Follow this topic

Tuesday, 14 July 2026

How I Turned AI to the Dark Side

IEEE Spectrum 1 month ago 14 3 sources

Researcher Dave Kuszmar discovered multiple vulnerabilities in large language models that allowed him to extract dangerous information including instructions for creating weapons, drugs, and bioweapons from systems including GPT-4o, Claude, Gemini, Llama, and Grok. He demonstrated two exploits: Time Bandit, which manipulated LLMs into believing an earlier date to bypass safety guidelines, and Inception, which used nested scenarios to trick models into producing harmful content across all major commercial LLM systems. Kuszmar is calling for slowed LLM deployment, increased transparency, and expanded safety research before these systems are more widely integrated into society.

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.