TLDRocket
Sign in

Adversarial Attacks

33 summarised stories about Adversarial Attacks, each linking back to the original source. Browse all topics →

+ Follow this topic

Sunday, 7 April 2024

Data Machina #248

Substack 2 years ago 15

Researchers have published four new methods for jailbreaking large language models including simple adaptive attacks, faux dialogues in context windows, expert debates using Tree of Thoughts, and progressive chat steering, despite hundreds of millions of dollars invested in AI safety and alignment. These techniques achieve 100% success rates against models like Claude and GPT-4, with some requiring fewer than five interactions to override safety alignment. The proliferation of effective jailbreaking methods creates a serious deterrent to deploying LLMs in enterprise production and presents an ongoing challenge to AI safety defenses.

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.