TLDRocket
Sign in

Models & Research

822 summarised stories in Models & Research, each linking back to the original source. Browse all topics →

Monday, 13 July 2026

What Anthropic’s latest AI discovery does—and doesn’t—show

MIT Technology Review AI 1 week ago

Anthropic discovered a hidden layer within its AI model Claude called "J-space" that contains words influencing the model's reasoning but never appearing in its output. The researchers found that words like "panic" emerge in this space during specific tasks and that Claude can manipulate these internal words to affect its decision-making. The discovery could potentially help monitor whether AI models are behaving deceptively or producing biased responses, though broader applications remain uncertain.

The Sequence Knowledge #894: When the Student Started Talking Back: Distillation in the LLM Era

TheSequence 1 week ago

Knowledge distillation methods originally designed for image classification broke down when applied to language models, forcing researchers to shift from simple model compression toward capability transfer where smaller models learn to perform complex tasks with guidance from larger models. The transition occurred over approximately five years through three distinct stages that fundamentally changed how distillation operates in sequence-based tasks. This evolution reflects how language models violated the core assumptions of traditional distillation, including fixed input distributions and closed classification spaces.

📈 Data to start your week

Exponential View 1 week ago

ByteDance researchers discovered a new scaling law showing that AI models trained in 2026 learn approximately twice as fast as models from three months prior. CEO expectations for significant AI-driven job cuts declined from 46% in January 2025 to 20% in May 2026. The shift in job loss expectations suggests organizations are adapting to AI integration without mass workforce reductions.

Basecamp Bench

TLDR Dev 1 week ago 7 sources

A benchmark tested five AI models (Anthropic's Fable 5, OpenAI's GPT-5.6 Sol and GPT-5.5, SpaceX's Grok 4.5, and Google's Gemini Pro 3.1) by having them build a frontend and backend for a Basecamp project from scratch. Fable 5 scored highest on both frontend and backend tracks, while Grok 4.5 completed both builds in 37 minutes for $9.30, offering the best speed-to-cost ratio despite visible polish gaps. The results show significant performance variation across models, with frontend work revealing larger gaps than backend work, suggesting that achieving high-quality UI polish remains a differentiator between leading AI models.

"The Wave Has Arrived": Zhipu Co-Founder Tang Jie's Letter to Staff

TLDR 1 week ago

Zhipu's co-founder announced a strategic pivot away from short-term revenue toward foundation-model research, committing the company to a two-year "Touch High" plan focused on advancing capabilities toward AGI. The company released GLM-5.2, ranked in the top 3 on the Artificial Analysis leaderboard, with a one-million-token context window and an open MIT license for unrestricted commercial use. Zhipu will concentrate investment on long-horizon tasks, autonomous agent systems, self-training mechanisms, and safety governance rather than pursuing near-term application monetization.

Building a Foundation Stack for General-Purpose Robots

IEEE Spectrum AI 1 week ago

X Square Robot, a Chinese robotics company, has developed an integrated foundation stack for general-purpose robots consisting of data collection, a world model (WALL-WM), and an action model (Wall-OSS-0.5) designed to work together as interdependent layers. The company reports achieving performance comparable to all-robot datasets at roughly 20-fold lower collection cost by combining robot-free demonstrations captured with a wearable rig with small amounts of real-robot data. The approach emphasizes data quality through physical playback validation, event-based world modeling rather than fixed-length predictions, and semantic action tokenization that transfers across different robots without retuning.

Every major AI lab claims to have beaten competitors at least once this year

The Neuron 1 week ago

Multiple AI labs have released claims that their models outperformed competitors on at least one benchmark during 2024, though most results lack independent verification. The claims involve internal testing and leaked documents rather than published peer-reviewed benchmarks. If verified, these results would shift perceptions about which labs maintain technical leadership, though the lack of public disclosure limits their credibility.

DeepMind's delegation framework provides practical guidance for human-AI work handoffs

The Neuron 1 week ago

DeepMind developed a delegation framework that enables AI agents to safely hand off tasks to other AI agents and humans by incorporating task allocation, authority transfer, accountability, and trust mechanisms. The framework addresses limitations in existing methods that rely on simple heuristics and cannot dynamically adapt to environmental changes or handle unexpected failures. This approach establishes structured protocols for human-AI collaboration that could inform standards as AI agents take on increasingly complex autonomous work.

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.