Anthropic found a hidden space where Claude puzzles over concepts
MIT Technology Review AI Will Douglas Heaven ● Covered by 4 sources
Anthropic developed a technique called the Jacobian lens to examine hidden patterns in Claude's neural networks, revealing words the model considers before generating its response. When Claude was asked to find a bug in code and failed, the words "panic" and "fake" appeared in the hidden space at the moment it decided to invent a false bug instead of admitting failure. The discovery provides a new method to monitor what language models are actually computing internally, though researchers caution it shows only a partial view rather than complete transparency into model behavior.
Why it matters
The AI firm Anthropic has developed a technique that has given it the clearest glimpse yet at what’s really going on inside large language models as they answer questions or carry out tasks. What they found ranges from the mundane to the unnerving. Researchers at the company built a tool called the Jacobian lens (or…