GPT-Red
GPT-Red is an internal automated red-teaming model developed by OpenAI using self-play reinforcement learning to identify vulnerabilities in its own models, particularly prompt injection attacks. In testing, GPT-Red achieved an 84% success rate on indirect prompt injection benchmarks against GPT-5.1 compared to 13% for human red-teamers, and discovered a novel attack class called Fake Chain-of-Thought. Training against GPT-Red's attacks resulted in a six-fold reduction in prompt injection failures on production models compared to earlier versions, shifting security testing from manual human discovery to continuous automated adversarial testing.
Updated 3 August 2026
Specifications
No specifications recorded yet.
Latest developments
OpenAI Details GPT-Red: An Internal Automated Red-Teaming Model That Beat Human Red-Teamers 84% To 13% On Prompt Injection
MarkTechPost · 2 weeks ago ·
23
Deep Learning Weekly: Issue 464
Deep Learning Weekly · 2 weeks ago ·
10
OpenAI’s GPT-Red automates prompt injection testing to harden AI agents
The New Stack · 2 weeks ago ·
28
Meet GPT-Red: an LLM super-hacker OpenAI built to make its models safer
CSET Georgetown · 2 weeks ago ·
45
Meet GPT-Red: an LLM super-hacker OpenAI built to make its models safer
MIT Technology Review AI · 2 weeks ago ·
31
GPT-Red: Unlocking Self-Improvement for Robustness
OpenAI Blog · 2 weeks ago ·
23
2026
Thinking Machines Lab releases Inkling, a 975B-parameter open-weights multimodal mixture-of-experts model Open source release
OpenAI releases GPT-Red, an automated red-teaming system for identifying vulnerabilities in language models Product launch
Researcher discovers vulnerabilities in major LLMs; OpenAI releases automated red teaming system Security issue
- OpenAI Details GPT-Red: An Internal Automated Red-Teaming Model That Beat Human Red-Teamers 84% To 13% On Prompt Injection
- Deep Learning Weekly: Issue 464
- OpenAI’s GPT-Red automates prompt injection testing to harden AI agents
- Meet GPT-Red: an LLM super-hacker OpenAI built to make its models safer
- Meet GPT-Red: an LLM super-hacker OpenAI built to make its models safer
- GPT-Red: Unlocking Self-Improvement for Robustness
Relationships
Products & technology
Competition
- Competes with GPT-5.1 · 1 source