Anthropic found Claude's hidden J-space for silent reasoning
X ● Covered by 4 sources
Anthropic found a hidden space inside Claude where it silently works out answers first. Shut it off and Claude still writes well but reasons far worse.
Based on reporting by X — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Anthropic's interpretability team has spotted something they're calling J-space, an internal holding area inside Claude where the model appears to stash, tweak, and reuse concepts before any of that thinking shows up in the text it produces. Think of it less as a scratchpad and more as a shared whiteboard that multiple parts of the network can read from and write to at once, pulling in relevant ideas mid-thought rather than starting from scratch every time.
The researchers found this by poking at Claude's internals directly rather than just reading its outputs. When they suppressed J-space activity, the model didn't fall apart. It kept producing grammatically clean, fluent sentences, the kind of output that looks perfectly normal on the surface. But its ability to work through anything genuinely complex — multi-step logic, harder reasoning tasks — dropped noticeably. The fluency stayed. The actual thinking didn't.
That gap is the interesting part. It suggests Claude's visible chain-of-thought, the step-by-step explanation it gives when asked to reason out loud, isn't the whole story of how it arrives at conclusions. There's apparently a layer of computation happening underneath that text, doing real work that never gets written down. For a field that's leaned hard on reading chain-of-thought as a window into model reasoning — and even proposed using it as a safety monitoring tool — that's a complication worth sitting with.
Anthropic has spent the past couple of years building out this kind of mechanistic interpretability work, mapping features and circuits inside its models the way earlier researchers mapped neurons in biological brains. J-space fits into that broader project: another piece of evidence that these systems have internal structure and internal processes that don't map cleanly onto the words they eventually type out. Whether that structure can be reliably read, audited, or trusted is still very much an open question, and this finding just made the inside of the black box a bit more complicated than the outside suggested.
My take — AI-written commentary, not fact-checked reporting
I keep hearing people treat chain-of-thought as basically a transcript of what a model is thinking, and this is exactly why that's naive — there's clearly a whole layer of computation happening off the page. If Anthropic is right that real reasoning can happen somewhere text never touches, every plan to use visible chain-of-thought as a safety check needs a rewrite, and the closed labs doing this kind of deep interpretability work are, for once, earning some of the trust they keep asking for.
Read more about this at: X