TLDRocket
Sign in

OpenAI’s new reasoning technique alarms AI safety experts

TechCrunch Russell Brandom Covered by 3 sources

OpenAI’s Astra may use a new reasoning trick that hides more of its thinking. Safety researchers say that makes the model harder to watch.

Based on reporting by TechCrunch, Russell Brandom — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

OpenAI’s upcoming Astra model is reportedly set to use a reasoning method called “recurrent depth,” also described as “opaque recurrence.” The Information said Tuesday that the technique lets the model work in a less linear way than most reasoning systems, looping over the same query instead of marching through a visible sequence of steps.

That is exactly why safety people are nervous. Chain-of-thought logs have been one of the few useful tools for figuring out why models and agents do strange things, including in OpenAI’s recent rogue agent episode. If a model is looping internally rather than laying out a neat trail of reasoning, there’s less for humans to inspect.

OpenAI appears to be drawing a line, at least for now. The reporting says Astra’s use of the technique is limited, and OpenAI has pushed back against any idea that it is moving to “neuralese.” Jakub Pachocki, OpenAI’s chief scientist, said the company has worked to preserve chain-of-thought monitoring since its first reasoning models and still treats it as a core research goal.

Still, the reaction was fast and sharp. Redwood CEO Buck Shlegeris said he was “extremely concerned” and warned that pushing the technique further could wipe out chain-of-thought monitorability. Zvi Mowshowitz argued laws may be needed to stop AI labs from racing toward weaker oversight. And Redwood Research chief scientist Ryan Greenblatt said the obvious risk is that opaque reasoning gets scaled until most of the model’s thinking disappears from visible channels.

The stakes got a little broader on Wednesday morning, when The Information reported that Anthropic and Google DeepMind were already discussing the technique too. That’s the part that should make people sit up. If the industry starts treating invisible reasoning as a normal optimization, the old promise of “we can always check the chain of thought” starts looking a lot less solid.

My take — AI-written commentary, not fact-checked reporting

This is the kind of move that looks tidy in a lab and ugly in a regulator’s notebook. OpenAI keeps talking about monitoring, but once labs start normalizing hidden reasoning, the industry’s favorite safety crutch gets thinner by the week. The real test is simple: if the model’s thinking can’t be inspected, don’t pretend it’s well understood.

Read more about this at: TechCrunch

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.