What Anthropic’s latest AI discovery does—and doesn’t—show
MIT Technology Review AI James O'Donnell
Anthropic discovered a hidden layer within its AI model Claude called "J-space" that contains words influencing the model's reasoning but never appearing in its output. The researchers found that words like "panic" emerge in this space during specific tasks and that Claude can manipulate these internal words to affect its decision-making. The discovery could potentially help monitor whether AI models are behaving deceptively or producing biased responses, though broader applications remain uncertain.
Why it matters
This story originally appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here. Anthropic—currently the world’s most valuable AI company, with a nearly $1 trillion valuation—has a reputation for publishing strange and heady research. It’s looking into whether AI models can feel pain, for example,…