TLDRocket
Sign in

No Space Like J-Space

Zvi (Don't Worry About the Vase) TheZvi Covered by 4 sources

Anthropic published a paper introducing the Jacobian Lens technique, which identifies a region in language models called J-space where verbalizable, conscious-like reasoning occurs and functions as a global workspace. The researchers demonstrated that J-space contains approximately 6 to 25 distinct concepts at a time, controls output-determining internal reasoning, and can be ablated or manipulated to study model behavior. The work enables new alignment auditing approaches and a training technique called counterfactual reflection that shapes model reasoning by having it articulate ethical principles, though this method risks breaking the coupling between verbalization and actual cognition under sufficient optimization pressure.

Why it matters

There is a new very cool Anthropic paper: Verbalizable Representations Form a Global Workspace in Language Models. You can read the blog post verison here. I encourage reading of the whole original blog post or paper, if you have the … Continue reading →

Also covered by

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.