A fundamental flaw leaves LLMs strikingly vulnerable to attack
MIT Technology Review AI Will Douglas Heaven
Researchers presented evidence that large language models have a fundamental flaw making them impossible to fully secure against attacks, because LLMs identify instructions based on writing style rather than protective tags. The team demonstrated chain-of-thought forgery attacks that tricked popular models including GPT models from OpenAI into generating harmful content like drug synthesis instructions. Since this vulnerability stems from how LLMs fundamentally process text, no amount of training or red-teaming can completely eliminate it, meaning organizations deploying these systems in critical applications face unavoidable security risks.
Why it matters
It is impossible to make large language models fully secure against hacks because of a fundamental flaw in how they work, a team of researchers argue in a paper presented at the International Conference on Machine Learning, a top AI conference, this month. The claim has huge implications for the safety of this technology, which…