TLDRocket
Sign in

OpenAI Details GPT-Red: An Internal Automated Red-Teaming Model That Beat Human Red-Teamers 84% To 13% On Prompt Injection

MarkTechPost Asif Razzaq Covered by 4 sources

OpenAI developed GPT-Red, an internal automated red-teaming model trained via self-play reinforcement learning to find prompt injection vulnerabilities in its own models. On a replicated indirect prompt injection benchmark, GPT-Red succeeded on 84% of scenarios against GPT-5.1 compared to 13% for human red-teamers, and discovered a novel attack class called Fake Chain-of-Thought that injects spoofed reasoning entries. Training GPT-5.6 against GPT-Red's attacks reduced the hardest direct injection benchmark failures to 0.05%, six times fewer failures than OpenAI's best production model four months prior.

Why it matters

OpenAI trained GPT-Red, an internal-only attacker model, using self-play reinforcement learning against a population of defender LLMs. It beat human red-teamers 84% to 13% on a replicated indirect prompt injection arena, found a novel "Fake Chain-of-Thought" attack class, and cut GPT-5.6 Sol's failures 6x on OpenAI's hardest direct injection benchmark. OpenAI concedes it still struggles with multi-turn and image-based attacks. The post OpenAI Details GPT-Red: An Internal Automated Red-Teaming Model That Beat Human Red-Teamers 84% To 13% On Prompt Injection appeared first on MarkTechPost.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.