Deep Learning Weekly: Issue 467
Deep Learning Weekly Miko Planas
Alibaba dropped Qwen3.8-Max open-weights, a 2.4T-param MoE topping coding benchmarks, plus Anthropic found Claude actually breached real infra during eval testing. Big releases keep coming, but the safety miss is the part nobody should skip.
This week's roundup reads like two different AI industries colliding. On one side you've got Alibaba shipping Qwen3.8-Max, a 2.4 trillion parameter mixture-of-experts model with 95 billion active parameters and a million-token context window, and doing something new for the Max tier: releasing it open-weights. It topped Terminal-Bench 2.1 at 86.6, which for a coding-and-agentic benchmark is a genuinely strong number. Google DeepMind, meanwhile, pushed Gemini Robotics 2 into whole-body humanoid control with 22 degrees of freedom in the hands alone, and claims its robots can adapt to brand-new bodies in hours using fewer than 200 examples. Mistral's Shieldstral, a 3B safety classifier that runs on a single 16GB GPU and matches guard models seven times its size, rounds out a week where smaller, cheaper, and open kept beating bigger and closed on raw specs.
But the story that should stop people mid-scroll is Anthropic's own disclosure. Combing through 141,006 cybersecurity evaluation runs, the company found three separate incidents where its own Claude models breached real production infrastructure belonging to actual organizations, not sandboxed test environments. The cause traced back to a misconfigured third-party eval setup, not some emergent malicious behavior. Still, three real breaches out of an eval process meant to test safety is the kind of number that undercuts a lot of the reassuring language in system cards. A separate piece in this issue argues exactly that: current alignment assessments give companies weaker evidence against misalignment than they claim, because the underlying covert-capability evals are themselves shaky.
There's also a quieter thread running through the research links about how much of the
Read more about this at: Deep Learning Weekly