TLDRocket
Sign in

Meet GPT-Red: an LLM super-hacker OpenAI built to make its models safer

MIT Technology Review AI Will Douglas Heaven Covered by 3 sources

OpenAI built GPT-Red, an AI system that automatically finds vulnerabilities in its models by simulating cyberattacks, particularly prompt injection attacks where hackers embed hidden instructions to manipulate LLM behavior. In testing against GPT-5.6, fewer than 23% of GPT-Red's strongest attacks succeeded compared to over 90% against the earlier GPT-5 released in August last year. The system supplements human red-teaming efforts and enables OpenAI to identify new attack types before deployment, though it remains limited in handling multi-turn conversations and image-based attacks.

Why it matters

OpenAI has built an LLM super-hacker called GPT-Red that it uses as a sparring partner to help its other models boost their defenses against cyberattacks. Last week the company released the latest version of its flagship LLM, GPT-5.6. OpenAI says that training it against GPT-Red made the model its most robust release yet. GPT-Red automates…

Also covered by

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.