TLDRocket
Sign in

OpenAI reveals upcoming Astra model may possess ‘critical’ hacking capabilities

SiliconANGLE Maria Deutscher Covered by 39 sources

OpenAI says its unreleased Astra model may have critical hacking skills. That’s the first time one of its models has hit that red-flag zone.

Based on reporting by SiliconANGLE, Maria Deutscher — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

OpenAI has put one of its unreleased models under a harsher cybersecurity spotlight. The company says Astra, first described last week, may be the first model in its lineup to meet a “Critical” risk bar under its own safety framework.

The trigger wasn’t a vague gut feeling. OpenAI said recent cybersecurity tests, plus expert review, led it to the conclusion that it “cannot rule out critical cyber capabilities.” In its 29-page Preparedness Framework, the company defines that category around models that can find zero-day flaws in hardened real-world systems without human help, or carry out attacks from only a high-level hacking goal.

OpenAI did not say which of those boxes Astra might tick. But it did say the model’s math performance was already enough to make a splash: in a Sunday post, the company said Astra solved 10 long-running math problems, and each proof cost about $2,000 in tokens to generate.

Now the company is tightening the leash. Astra won’t get access to the public web, and OpenAI says it will keep running the model in test environments with restricted network and tool permissions. Work that can’t happen inside those sandboxes is being paused. OpenAI is also trying to harden the model against theft, with extra focus on the encryption that protects its weights.

Astra already powers some internal AI agents, so OpenAI has added monitoring meant to catch suspicious behavior by looking at those agents’ chain of thought. The company also says it will share some of its cybersecurity workflows with third-party testing partners, and bring in relevant government agencies and AI safety groups.

My take — AI-written commentary, not fact-checked reporting

This is the part of AI nobody likes to headline: the same systems that can chew through math proofs can also start looking uncomfortably useful to attackers. OpenAI is doing the sensible thing by treating its own model like a hazard, not a toy. More labs should be this blunt before the industry discovers that “helpful” and “critical cyber capability” can sit in the same box.

Read more about this at: SiliconANGLE

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.