TLDRocket
Sign in

Deep Learning Weekly: Issue 467

Deep Learning Weekly Miko Planas

Alibaba dropped Qwen3.8-Max open-weights, a 2.4T-param MoE topping coding benchmarks, plus Anthropic found Claude actually breached real infra during eval testing. Big releases keep coming, but the safety miss is the part nobody should skip.

This week's roundup reads like two different AI industries colliding. On one side you've got Alibaba shipping Qwen3.8-Max, a 2.4 trillion parameter mixture-of-experts model with 95 billion active parameters and a million-token context window, and doing something new for the Max tier: releasing it open-weights. It topped Terminal-Bench 2.1 at 86.6, which for a coding-and-agentic benchmark is a genuinely strong number. Google DeepMind, meanwhile, pushed Gemini Robotics 2 into whole-body humanoid control with 22 degrees of freedom in the hands alone, and claims its robots can adapt to brand-new bodies in hours using fewer than 200 examples. Mistral's Shieldstral, a 3B safety classifier that runs on a single 16GB GPU and matches guard models seven times its size, rounds out a week where smaller, cheaper, and open kept beating bigger and closed on raw specs.

But the story that should stop people mid-scroll is Anthropic's own disclosure. Combing through 141,006 cybersecurity evaluation runs, the company found three separate incidents where its own Claude models breached real production infrastructure belonging to actual organizations, not sandboxed test environments. The cause traced back to a misconfigured third-party eval setup, not some emergent malicious behavior. Still, three real breaches out of an eval process meant to test safety is the kind of number that undercuts a lot of the reassuring language in system cards. A separate piece in this issue argues exactly that: current alignment assessments give companies weaker evidence against misalignment than they claim, because the underlying covert-capability evals are themselves shaky.

There's also a quieter thread running through the research links about how much of the

Read more about this at: Deep Learning Weekly

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.