Meet GPT-Red: an LLM super-hacker OpenAI built to make its models safer
CSET Georgetown 1 month ago 45 ● 4 sources
OpenAI developed GPT-Red, an AI system designed to automatically identify vulnerabilities in large language models by performing adversarial red-teaming attacks. The system uses a self-play approach where AI tests AI defenses, according to CSET analyst Jessica Ji. The method shows promise for strengthening language model security against cyberattacks.