TLDRocket
Sign in

Safety & Ethics

428 summarised stories in Safety & Ethics, each linking back to the original source. Browse all topics →

Tuesday, 16 December 2025

Gemma Scope 2: helping the AI safety community deepen understanding of complex language model behavior

Google DeepMind 7 months ago

Google released Gemma Scope 2, an open-source suite of interpretability tools designed to help researchers understand the internal workings of its Gemma 3 language models. The release includes sparse autoencoders and transcoders trained on every layer of models ranging from 270 million to 27 billion parameters, built using approximately 110 petabytes of data. The tools enable AI safety researchers to debug emergent behaviors, audit AI agents, and develop defenses against jailbreaks, hallucinations, and other failure modes.

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.