TLDRocket
Sign in

The Specter Of Neuralese

Astral Codex Ten

OpenAI says its new looping model isn’t really the scary kind of AI “neuralese.” The fight is whether that distinction is real, or just a nicer label for the same risk.

Based on reporting by Astral Codex Ten — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

After the release of AI 2027, some critics said the report had overstated a danger: “neuralese recurrence,” the idea that an AI could keep thinking in a hidden internal language instead of in readable text. TLDR’s answer is that the worry is not imaginary, but the details matter a lot. OpenAI’s chief scientist, Jakub Pachocki, says the company’s system is “not really” doing the fully opaque version people fear.

The core issue is simple enough. Standard transformers do their thinking in layers, then often spill intermediate reasoning onto a chain-of-thought scratchpad in plain language before going back for another pass. That’s useful for capabilities, but also a gift for safety, because humans — or AI monitors — can read the scratchpad and catch bad intent. Neuralese recurrence would skip that readable step and keep the model’s intermediate thoughts in vector form, where the numbers are not easy for people to inspect.

Pachocki’s defense is that OpenAI is not jumping into that world. He says the depth of the computation graph for Astra is within a factor of two of GPT-4. The argument is that this is basically more depth, not some qualitatively new monster. A looped model can reuse layers to simulate a deeper network, but that is not the same as having the model think forever without ever surfacing anything in English.

And that distinction is doing a lot of work. A loop can help, and it can give the model more room to work through a problem, but it is not a free lunch. The piece points out that real layers add more than extra time: they can add new “room” for knowledge and styles of thought. Looping the same layers again and again is closer to asking the same mathematician to stare at the problem longer than to hiring a smarter mathematician.

That is why the safety debate gets slippery. If the model can do 1,000 unmonitored steps either by replacing all the blue, readable chain-of-thought with internal vector steps or by making a single forward pass absurdly deep between readable tokens, the danger starts to look similar from the outside. The article’s bottom line is that AI 2027’s truly opaque neuralese still does not exist in practice, but looped transformers move in that direction enough to keep people nervous.

My take — AI-written commentary, not fact-checked reporting

This is exactly why “we’re not really doing that” is not the comfort OpenAI thinks it is. If the safety line only survives when you squint at architecture diagrams, then the line is already wobbling. The EU would call this a governance problem; the rest of us can just call it the usual industry habit of renaming the cliff edge.

Read more about this at: Astral Codex Ten

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.