TLDRocket
Sign in

Deep Learning Weekly: Issue 450

Deep Learning Weekly Miko Planas

Google dropped Gemma 4, and a 31B open model is now punching above weight classes 20x its size. Meanwhile Anthropic and Alibaba are racing on agent infrastructure.

Based on reporting by Deep Learning Weekly, Miko Planas — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Deep Learning Weekly's latest issue reads like a snapshot of an industry that's stopped arguing about whether agents are the future and started arguing about who builds the plumbing for them fastest. Google's Gemma 4 is the headline open-weights story: a four-model family where the 31B parameter version lands at #3 among open models on Arena AI, beating out models twenty times its size. That's not a small claim, and it says something about where the gains are actually coming from these days — architecture and training recipes, not just raw parameter count.

But the more interesting fight might be happening one layer down, in the agent tooling itself. Anthropic pushed Claude Managed Agents into public beta, a set of cloud-hosted APIs meant to strip away the grunt work of running agents in production: sandboxing, state, permissions, orchestration. The pitch is blunt — ship in days, not months. Alibaba, not to be outdone, launched Qwen3.6-Plus, an agentic coding model claiming parity with or an edge over Claude Opus 4.5 on SWE-bench and Terminal-Bench 2.0. Two very different companies, same bet: the value isn't just in the model anymore, it's in making that model reliably do things in the real world.

Sebastian Raschka's piece on the components of a coding agent is worth lingering on, because it explains why tools like Claude Code or Codex CLI feel so different from just chatting with an LLM. He breaks it down into six architectural pieces that turn a raw model into something that can actually navigate a codebase and get work done — and it's a useful antidote to the marketing tendency to treat

My take — AI-written commentary, not fact-checked reporting

placeholder

Read more about this at: Deep Learning Weekly

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.