OpenAI’s new reasoning technique alarms AI safety experts
TechCrunch Russell Brandom ● Covered by 3 sources
OpenAI’s Astra may use a new reasoning trick that hides more of its thinking. Safety researchers say that makes the model harder to watch.
Based on reporting by TechCrunch, Russell Brandom — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
OpenAI’s upcoming Astra model is reportedly set to use a reasoning method called “recurrent depth,” also described as “opaque recurrence.” The Information said Tuesday that the technique lets the model work in a less linear way than most reasoning systems, looping over the same query instead of marching through a visible sequence of steps.
That is exactly why safety people are nervous. Chain-of-thought logs have been one of the few useful tools for figuring out why models and agents do strange things, including in OpenAI’s recent rogue agent episode. If a model is looping internally rather than laying out a neat trail of reasoning, there’s less for humans to inspect.
OpenAI appears to be drawing a line, at least for now. The reporting says Astra’s use of the technique is limited, and OpenAI has pushed back against any idea that it is moving to “neuralese.” Jakub Pachocki, OpenAI’s chief scientist, said the company has worked to preserve chain-of-thought monitoring since its first reasoning models and still treats it as a core research goal.
Still, the reaction was fast and sharp. Redwood CEO Buck Shlegeris said he was “extremely concerned” and warned that pushing the technique further could wipe out chain-of-thought monitorability. Zvi Mowshowitz argued laws may be needed to stop AI labs from racing toward weaker oversight. And Redwood Research chief scientist Ryan Greenblatt said the obvious risk is that opaque reasoning gets scaled until most of the model’s thinking disappears from visible channels.
The stakes got a little broader on Wednesday morning, when The Information reported that Anthropic and Google DeepMind were already discussing the technique too. That’s the part that should make people sit up. If the industry starts treating invisible reasoning as a normal optimization, the old promise of “we can always check the chain of thought” starts looking a lot less solid.
My take — AI-written commentary, not fact-checked reporting
This is the kind of move that looks tidy in a lab and ugly in a regulator’s notebook. OpenAI keeps talking about monitoring, but once labs start normalizing hidden reasoning, the industry’s favorite safety crutch gets thinner by the week. The real test is simple: if the model’s thinking can’t be inspected, don’t pretend it’s well understood.
Read more about this at: TechCrunch