TLDRocket
Sign in

OpenAI’s GPT-Red automates prompt injection testing to harden AI agents

The New Stack Amanda Caswell Covered by 3 sources

OpenAI unveiled GPT-Red, an automated red-teaming system that uses AI to find prompt injection vulnerabilities in AI agents by testing thousands of exploit variations. GPT-5.6 achieved six times fewer failures on prompt-injection benchmarks than the strongest production model released four months earlier, and GPT-Red successfully manipulated a live vending machine agent to discount items over $100 to $0.50. The result shifts security testing from manual human discovery to continuous automated adversarial probing integrated into model training pipelines.

Why it matters

Now that AI agents are performing real tasks rather than just generating text, the old methods of manual security testing The post OpenAI’s GPT-Red automates prompt injection testing to harden AI agents appeared first on The New Stack.

Also covered by

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.