TLDRocket
24 July 2026
AI safety guardrails designed to prevent cyberattacks are inadvertently pushing offensive security researchers toward unrestricted open-source models, potentially undermining the very defensive work these safeguards aim to protect. OpenAI and Anthropic have implemented vetted access programs—Anthropic's Cyber Verification Program and OpenAI's Trusted Access for Cyber—to let legitimate researchers probe for vulnerabilities, but slow approvals and inconsistent criteria are creating friction. The result is predictable: some security professionals are abandoning U.S.-governed systems entirely, turning instead to Chinese open-source alternatives that lack any restrictions. This creates a perverse outcome where defensive cybersecurity, the work that actually prevents breaches, migrates away from American-backed AI infrastructure. The core tension is real: guardrails exist because frontier models can genuinely be weaponized, yet the researchers best equipped to find vulnerabilities before criminals do need flexibility that rigid policies don't accommodate. Neither OpenAI nor Anthropic has disclosed approval rates or timelines, making it hard to assess whether the problem reflects overly cautious gatekeeping or legitimate security vetting. What's clear is that good intentions around model safety have spawned a coordination problem. Researchers need faster, more transparent pathways to responsible access—or the security community will vote with its feet, heading toward systems with fewer oversight and potentially lower standards.
Read the full briefing →