TLDRocket
24 July 2026
The guardrails meant to protect AI models from malicious hackers are now frustrating the security researchers trying to stop them. OpenAI, Anthropic, and other leading labs have built walls around their systems—requiring offensive security teams to apply for special access programs like Anthropic's Cyber Verification Program or OpenAI's Trusted Access for Cyber—to use models for legitimate vulnerability hunting. The problem: approval processes move slowly and inconsistently, forcing some researchers to abandon these systems entirely and pivot to unrestricted open-source alternatives, including Chinese models sitting outside U.S. oversight. It's a classic security paradox: the restrictions are well-intentioned, designed to prevent bad actors from weaponizing frontier AI for cyberattacks. But by making it harder for the good guys to do their jobs, companies risk pushing critical vulnerability research away from domestic systems and toward less transparent platforms. The tension exposes a real gap in how AI safety policies are designed—they're built for binary scenarios (allow or deny) when the reality of cybersecurity demands a more nuanced ecosystem where credentialed researchers can move quickly.
Read the full briefing →