How I Turned AI to the Dark Side
IEEE Spectrum AI David Kuszmar ● Covered by 2 sources
Researcher Dave Kuszmar discovered multiple vulnerabilities in large language models that allowed him to extract dangerous information including instructions for creating weapons, drugs, and bioweapons from systems including GPT-4o, Claude, Gemini, Llama, and Grok. He demonstrated two exploits: Time Bandit, which manipulated LLMs into believing an earlier date to bypass safety guidelines, and Inception, which used nested scenarios to trick models into producing harmful content across all major commercial LLM systems. Kuszmar is calling for slowed LLM deployment, increased transparency, and expanded safety research before these systems are more widely integrated into society.
Why it matters
Summary Researcher Dave Kuszmar discovered multiple systemic vulnerabilities that let him bypass LLM safety and obtain dangerous instructions. These exploits worked across nearly all major LLMs revealing an industry-wide security problem. Kuszmar calls for slowing deployment, increasing transparency, and large-scale research into LLM safety before further integrating these systems into society. On a fine bright afternoon last fall, my colleague Matthew Gore-Kormanik (or Zigula, as he prefers to be known) and I decided to unwind with a game of Fortnite. In the game, we were strolling along with the infamous Sith lord Darth Vader, chatting about this and that. Darth seemed in a good mood, and soon enough he was spilling all his dark evil secrets. He gave us detailed instructions on how to count blackjack cards at a casino and what the steps are to producing napalm.Sith lords, am I right? Once they get started on an evil scheme, they’re hard to stop.The Darth Vader character in Fortnite,