TLDRocket
Sign in

OpenAI Details GPT-Red: An Internal Automated Red-Teaming Model That Beat Human Red-Teamers 84% To 13% On Prompt Injection

MarkTechPost Asif Razzaq Covered by 3 sources

OpenAI developed GPT-Red, an internal automated red-teaming model trained via self-play reinforcement learning to find prompt injection vulnerabilities in its own models. On a replicated indirect prompt injection benchmark, GPT-Red succeeded on 84% of scenarios against GPT-5.1 compared to 13% for human red-teamers, and discovered a novel attack class called Fake Chain-of-Thought that injects spoofed reasoning entries. Training GPT-5.6 against GPT-Red's attacks reduced the hardest direct injection benchmark failures to 0.05%, six times fewer failures than OpenAI's best production model four months prior.

Why it matters

OpenAI trained GPT-Red, an internal-only attacker model, using self-play reinforcement learning against a population of defender LLMs. It beat human red-teamers 84% to 13% on a replicated indirect prompt injection arena, found a novel "Fake Chain-of-Thought" attack class, and cut GPT-5.6 Sol's failures 6x on OpenAI's hardest direct injection benchmark. OpenAI concedes it still struggles with multi-turn and image-based attacks. The post OpenAI Details GPT-Red: An Internal Automated Red-Teaming Model That Beat Human Red-Teamers 84% To 13% On Prompt Injection appeared first on MarkTechPost.

Also covered by

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.