TLDRocket
Sign in

GPT-Red

Model Covered in 6 stories Compare ⇄ + Follow

GPT-Red is an internal automated red-teaming model developed by OpenAI using self-play reinforcement learning to identify vulnerabilities in its own models, particularly prompt injection attacks. In testing, GPT-Red achieved an 84% success rate on indirect prompt injection benchmarks against GPT-5.1 compared to 13% for human red-teamers, and discovered a novel attack class called Fake Chain-of-Thought. Training against GPT-Red's attacks resulted in a six-fold reduction in prompt injection failures on production models compared to earlier versions, shifting security testing from manual human discovery to continuous automated adversarial testing.

Updated 3 August 2026

Specifications

No specifications recorded yet.

Latest developments

Timeline

Month Quarter Year

2026

Thinking Machines Lab releases Inkling, a 975B-parameter open-weights multimodal mixture-of-experts model Open source release

OpenAI releases GPT-Red, an automated red-teaming system for identifying vulnerabilities in language models Product launch

Researcher discovers vulnerabilities in major LLMs; OpenAI releases automated red teaming system Security issue

Relationships

Products & technology

  • OpenAI develops this model · 6 sources
  • Integrated with GPT-5.6 · 1 source

Competition

  • Competes with GPT-5.1 · 1 source

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.