TLDRocket
Sign in

Claude’s cyber safeguards are getting more flexible — but not for everyone.

The New Stack Meredith Shubel ● Covered by 2 sources

Anthropic split its cyber access into three tiers, and most teams still get the locked-down one. Its own tests show the strictest path blocked Claude Opus 5.5 in 46 of 50 offensive trials.

Based on reporting by The New Stack, Meredith Shubel — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Anthropic is loosening its Cyber Verification Program, but only in a very controlled way. The company has split access into three use case-based tiers and moved Project Glasswing into the least restricted one. On paper, that broadens access for cyber defenders. In practice, Anthropic’s own results say the guardrails still bite hard.

The default for many defensive users will be Defense Access. Anthropic says that tier is meant for incident response and vulnerability analysis, and it is the broadest in who can apply, covering smaller security firms, open-source maintainers, and individual researchers with a record of reported vulnerabilities. But the company’s test results show Claude Opus 5.5 under those safeguards was blocked in 46 of 50 trials on CyScenario Bench, Anthropic’s evaluation for planning and carrying out multi-stage cyber operations. Without CVP access at all, every task was stopped on the first prompt.

That doesn’t mean the tier is failing at its job. Anthropic says those tests are “complex, interactive offensive scenarios,” so it expected major blocks. Still, the company doesn’t yet have a comparable defensive benchmark, which leaves a real blind spot: defenders can’t tell how much the controls will get in the way of legitimate work.

The middle tier, Red Team Access, gives much more room. Anthropic says it is aimed at red teams and penetration-testing firms, while individual researchers are barred for now. Under those safeguards, Opus 5.5 completed 34 of 50 tasks, and none were blocked. The top tier, Specialized Access, is reserved for a limited set of verified organizations testing safety-critical systems like power grids or telecom networks. Project Glasswing members keep that access, but new applicants need a review from Anthropic and the U.S. government.

There is also a data catch. Everyone in CVP has to accept data retention, at least for now. Anthropic says Enterprise Frontier Safeguards is coming later this fall for eligible organizations that want self-controlled cloud storage, and users with zero-retention access to Claude Fable 5.1 or Claude Mythos 5.1 are exempt in the meantime. The company also says it is still working on ways to help secure open-source software and critical infrastructure, but details are coming later.

My take — AI-written commentary, not fact-checked reporting

This is Anthropic doing the classic safety-company move: open the door, then bolt a second lock on the frame. The split tiers make sense if the goal is real control, not a marketing brochure with a padlock emoji. What’s missing is a serious defensive benchmark; without that, “more flexible” mostly means “more paperwork.”

Read more about this at: The New Stack

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.