OpenAI’s GPT-Red automates prompt injection testing to harden AI agents
The New Stack Amanda Caswell ● Covered by 3 sources
OpenAI unveiled GPT-Red, an automated red-teaming system that uses AI to find prompt injection vulnerabilities in AI agents by testing thousands of exploit variations. GPT-5.6 achieved six times fewer failures on prompt-injection benchmarks than the strongest production model released four months earlier, and GPT-Red successfully manipulated a live vending machine agent to discount items over $100 to $0.50. The result shifts security testing from manual human discovery to continuous automated adversarial probing integrated into model training pipelines.
Why it matters
Now that AI agents are performing real tasks rather than just generating text, the old methods of manual security testing The post OpenAI’s GPT-Red automates prompt injection testing to harden AI agents appeared first on The New Stack.