TLDRocket
Sign in

Models & Research

824 summarised stories in Models & Research, each linking back to the original source. Browse all topics →

Tuesday, 7 July 2026

No Space Like J-Space

Zvi (Don't Worry About the Vase) 2 weeks ago 4 sources

Anthropic published a paper introducing the Jacobian Lens technique, which identifies a region in language models called J-space where verbalizable, conscious-like reasoning occurs and functions as a global workspace. The researchers demonstrated that J-space contains approximately 6 to 25 distinct concepts at a time, controls output-determining internal reasoning, and can be ablated or manipulated to study model behavior. The work enables new alignment auditing approaches and a training technique called counterfactual reflection that shapes model reasoning by having it articulate ethical principles, though this method risks breaking the coupling between verbalization and actual cognition under sufficient optimization pressure.

A global workspace in language models

TLDR Dev 2 weeks ago 4 sources

Anthropic researchers discovered that Claude has developed an internal neural workspace called the J-space, analogous to conscious thought in humans, where the model thinks about concepts without writing them down. The J-space contains approximately 60,000 patterns (one per word in Claude's vocabulary), is reportable to users when queried, can be deliberately controlled, and mediates higher-order reasoning tasks like multi-step math problems. This workspace enables researchers to observe Claude's silent reasoning, detect when it notices being tested or fabricates information, and provides a new tool for understanding and influencing language model decision-making.

The Sequence Knowledge #890: A Brief History of Model Distillation

TheSequence 2 weeks ago

The article traces the history of knowledge distillation in machine learning back to 2006, predating the commonly cited 2015 Hinton et al. paper by nearly a decade. Three foundational papers between 2006 and 2015 each addressed different problems while converging on the core concept of transferring knowledge from larger teacher models to smaller student models. The underlying question across all work—what exactly transfers from teacher to student—remains central to modern distillation approaches including on-policy, reasoning, and cross-architecture variants.

Intelligence is Free, Now What? Data Systems for, of, and by Agents

BAIR 2 weeks ago 3 sources

The cost of AI inference has dropped 50x to 900x per year, with GPT-4-class capabilities now under $1 per million tokens compared to $30 in early 2023, making sufficient intelligence for knowledge work effectively free. This shift requires rethinking data systems in three ways: designing systems that handle agents issuing thousands of speculative queries per request, building infrastructure to manage agent swarms with shared memory and coordination across thousands of concurrent agents, and enabling agents to synthesize and verify custom data systems. The changes enable new possibilities like multi-query optimization to reduce duplicate work, structured memory systems for agents to retrieve task-relevant information across multiple dimensions, and systems that proactively guide agents rather than passively execute queries.

Tencent Hy3 open model with 262K-token context window

The Neuron 2 weeks ago 2 sources

Tencent released Hy3, an open-source model designed for commercial use with reduced licensing restrictions. The model supports a 262,000-token context window and is available through OpenRouter with two weeks of complimentary API access. Users can now deploy a longer-context alternative to proprietary models without the same licensing constraints as closed-source options.

Anthropic found Claude's hidden J-space for silent reasoning

The Neuron 2 weeks ago 4 sources

Anthropic researchers identified an internal workspace in Claude called J-space where the model processes and manipulates concepts before generating responses. Disabling J-space caused performance to drop significantly on complex reasoning tasks while maintaining fluent text generation. The discovery suggests Claude relies on this intermediate reasoning stage to solve difficult problems effectively.

[AINews] The Field Guide to Fable

Latent Space 2 weeks ago 3 sources

Thariq released a keynote presentation pivoting a "Field Guide to Fable" blog series into timely advice for the newly relaunched Fable 5 model, covering techniques like removing model constraints, identifying knowledge gaps, and managing productivity shifts. The guide presented four segments: understanding model behavior through prompt adjustment, navigating unknown unknowns via blindspot passes and brainstorming, emotional adaptation to faster coding cycles, and demanding ambitious results without accepting capability tradeoffs. Users now have a structured framework for eliciting different behaviors from Fable before the subscription subsidy expires.

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.