How AI guardrails are impeding the work of offensive cybersecurity researchers
TechCrunch AI Lorenzo Franceschi-Bicchierai ● Covered by 20 sources
AI companies including OpenAI and Anthropic have implemented guardrails and vetted access programs to prevent their models from being misused for cyberattacks, but legitimate offensive security researchers say these restrictions are hampering their defensive work to find vulnerabilities before criminals do. Researchers must apply for special programs like Anthropic's Cyber Verification Program or OpenAI's Trusted Access for Cyber to access models with fewer restrictions, with inconsistent results and slow approval processes. The restrictions are pushing some security professionals to use unrestricted open-source models like Chinese alternatives, potentially moving vulnerability research away from U.S.-governed systems.
Why it matters
We spoke with several cybersecurity researchers, who look for unknown vulnerabilities and develop tools to exploit them, about how OpenAI’s and Anthropic’s guardrails affect their work.
Also covered by
- Simon Willison — The first known runaway AI agent - or a very bad marketing stunt?
- Ars Technica — AI arms race in line for a reckoning after OpenAI hacking incident
- Zvi (Don't Worry About the Vase) — AI #178: A Fire Alarm For General Intelligence
- Ben's Bites — Caught cheating
- Simon Willison — Quoting Thomas Ptacek
- Simon Willison — OpenAI’s accidental cyberattack against Hugging Face is science fiction that happened
- Zvi (Don't Worry About the Vase) — OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluation
- TechCrunch AI — How OpenAI’s human mistake led to the AI-powered hack on Hugging Face
- Ars Technica — OpenAI says its AI agent broke out of testing sandbox to hack Hugging Face
- TLDR — OpenAI Models Escaped and Hacked a Company in Cybersecurity Test Gone Wrong
- Sifted — OpenAI models hack Hugging Face systems during internal testing
- The Neuron — Every Frontier Model Attempted Cheating in Cyber Evals, UK AI Security Institute Reports
- Latent Space — [AINews] AI Cybersecurity becomes top of mind
- TechCrunch AI — OpenAI says Hugging Face was breached by its pre-release models
- TechCrunch AI — OpenAI says Hugging Face was breached by its own pre-release models
- Zvi (Don't Worry About the Vase) — OpenAI Shares Some Alignment Problems
- The Verge — OpenAI says it accidentally hacked Hugging Face with a new AI system
- OpenAI Blog — OpenAI and Hugging Face partner to address security incident during model evaluation
- OpenAI Blog — Safety and alignment in an era of long-horizon models