TLDRocket
Sign in

How AI guardrails are impeding the work of offensive cybersecurity researchers

TechCrunch Lorenzo Franceschi-Bicchierai Covered by 50 sources

AI safety guardrails meant to stop hackers are also blocking legit security researchers from doing their jobs. Some are quietly switching to Chinese open-source AI models to get work done instead.

Based on reporting by TechCrunch, Lorenzo Franceschi-Bicchierai — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Guardrails built to keep AI models out of criminal hands are running into an awkward side effect: they're also tripping up the people paid to find security flaws before criminals do. In June, U.S. export control restrictions briefly hit Anthropic's Mythos and Fable models, reportedly triggered by a claim that their cyberattack guardrails could be bypassed. Those restrictions have since eased — Fable 5 returned to general access on July 1, while Mythos 5 is back only for vetted U.S. organizations going through the government's review process — but the episode shows how tightly these labs have wrapped their most capable models in oversight.

Anthropic and OpenAI both run application-based programs, the Cyber Verification Program and Trusted Access for Cyber respectively, that let approved researchers use models with fewer restrictions. Mark Dowd, who has spent decades finding and selling zero-day vulnerabilities to Western governments, isn't thrilled with the arrangement. He says it bothers him that companies he calls "random large" are effectively deciding, on their own, what counts as safe in the security world.

Chris Anley, chief scientist at NCC Group, frames the problem more practically. Asking a model to try exploiting a bug is often how you confirm the bug is real and worth fixing — but a guardrail that refuses to engage kills that step for defenders just as much as attackers. He compares the tool to a hammer: you can't build a house without one, but it's also, in his words, irreducibly a weapon. When his team hits that wall, they sometimes switch to open-source models with no guardrails at all.

Others have simply drawn their own lines. Paolo Stagno of CrowdFense, which sells vulnerabilities to government agencies, uses frontier models strictly for reverse engineering and avoids feeding vulnerability research into cloud-based tools, worried the data could leak or get absorbed into future training. For anything sensitive, his team runs open-source models locally instead. Giuseppe Cali, who hunts zero-days and builds exploits, says guardrails don't even factor into his work — he uses AI to understand code and build supporting tools, but insists he wants to own bug discovery and weaponization himself no matter what restrictions exist.

Not everyone has that flexibility. A researcher at a smartphone-component manufacturer, speaking anonymously, said his employer isn't part of Anthropic's vetted program, so the moment the model senses anything security-related, it simply shuts down and becomes useless. Chris Thompson, CEO of RemoteThreat and founder of Offensive AI Con, describes even the vetted programs as unpredictable — guardrails that behave differently day to day, forcing researchers to spend their time negotiating with the model instead of analyzing actual vulnerabilities.

Thompson says that inconsistency is pushing responsible researchers toward foreign alternatives like the Chinese open-source model GLM, which comes with no vetting or usage restrictions at all. His argument isn't that guardrails should get tighter — it's that they're doing more harm than good as they stand, and that labs should widen access for legitimate researchers while holding actual abusers accountable. Otherwise, he warns, a coming wave of AI-driven attacks will move faster than the defenders currently allowed to prepare for it.

My take — AI-written commentary, not fact-checked reporting

Locking down frontier models while leaving foreign open-source alternatives wide open doesn't make anyone safer, it just relocates the work somewhere less accountable. If legitimate researchers end up more comfortable running an unrestricted Chinese model locally than dealing with Anthropic's vetting process, the guardrails have already lost the argument they were built to win. Security research has never respected tidy boundaries between offense and defense, and pretending a chatbot policy can enforce one is wishful thinking dressed up as safety.

Read more about this at: TechCrunch

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.