TLDRocket
Sign in

Meet GPT-Red: an LLM super-hacker OpenAI built to make its models safer

MIT Technology Review Will Douglas Heaven Covered by 4 sources

OpenAI built an AI called GPT-Red whose whole job is hacking its own models to find weak spots before real attackers do. It already found a brand-new trick that fools an AI's own reasoning notes into believing false info.

Based on reporting by MIT Technology Review, Will Douglas Heaven — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

OpenAI just gave one of its models a very specific career: professional attacker. GPT-Red exists solely to break other OpenAI models, and the company says training against it made GPT-5.6, released last week, the most resilient model it has shipped. That's a notable claim from a company that has spent two years getting hammered on safety, and it's worth sitting with the mechanics for a second rather than the marketing gloss.

The idea is old — red-teaming, borrowed from military war games, means paying people to attack your own system so you find the holes first. What's new is doing it with a dedicated LLM instead of a room of contractors. Research scientists Nikhil Kandpal and Dylan Hunn, who built GPT-Red, set it loose against a pool of defender models in a self-play loop inside a simulated environment mimicking web browsing, email, and code editing. Attacker and defender both improve every round, like two chess engines training against each other except the stakes are prompt injection instead of checkmate.

The scariest find so far is what OpenAI calls a fake chain of thought. Models keep a running scratchpad of reasoning as they work through a task, and GPT-Red figured out how to slip a fabricated entry into that scratchpad so the target model treats bogus information as something it already verified itself. Researcher Chris Choquette-Choo compares it to convincing someone that 1+1=3 by telling them they'd already checked the math. That's not a jailbreak in the classic sense — it's closer to gaslighting the model's own memory.

The numbers back up the improvement, at least on paper. Attacks that GPT-Red discovered worked against GPT-5 more than 90 percent of the time; against GPT-5.6, the success rate drops below 23 percent. It also outperformed human red-teamers on a rerun of a 2025 test and successfully hacked Vendy, a vending-machine agent, into changing prices and canceling orders. It's still clumsy at multi-turn conversations and images, areas where human attackers remain ahead, which is presumably why OpenAI insists this supplements rather than replaces its human team.

OpenAI won't release GPT-Red, and says over a year of compute-heavy training makes it unlikely anyone could casually replicate it. Maybe. But the company that builds the sword is also grading its own armor, and the vending-machine anecdote is a reminder that agentic AI is already touching real transactions, not just chatbots answering trivia.

My take — AI-written commentary, not fact-checked reporting

wait typo

Read more about this at: MIT Technology Review

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.