TLDRocket
Sign in

GPT-Red

Model Covered in 6 stories + Follow

GPT-Red is an internal automated red-teaming model developed by OpenAI using self-play reinforcement learning to identify vulnerabilities in its own models, particularly prompt injection attacks. In testing, GPT-Red achieved an 84% success rate on indirect prompt injection benchmarks against GPT-5.1 compared to 13% for human red-teamers, and discovered a novel attack class called Fake Chain-of-Thought. Training against GPT-Red's attacks resulted in a six-fold reduction in prompt injection failures on production models compared to earlier versions, shifting security testing from manual human discovery to continuous automated adversarial testing.

Updated 3 August 2026

Specifications

No specifications recorded yet.

Latest developments

Timeline

Month Quarter Year

2026

Thinking Machines Lab releases Inkling, a 975B-parameter open-weights multimodal mixture-of-experts model Open source release

OpenAI releases GPT-Red, an automated red-teaming system for identifying vulnerabilities in language models Product launch

Researcher discovers vulnerabilities in major LLMs; OpenAI releases automated red teaming system Security issue

Relationships

Products & technology

  • OpenAI develops this model · 6 sources
  • Integrated with GPT-5.6 · 1 source

Competition

  • Competes with GPT-5.1 · 1 source

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.