TLDRocket
Sign in

GLM 5.1 Thinks Strategically, Data-Center Revolt Intensifies, When Helpful LLMs Turn Unhelpful, Humanoid Robots Get to Work

The Batch Analytics DeepLearning.AI Covered by 2 sources

Z.ai's new GLM-5.1 can grind on a coding task for up to eight hours, replanning as it goes instead of quitting early. Meanwhile pushback against data centers is turning ugly, with two recent violent incidents alongside legislative moves like Maine's proposed moratorium.

Based on reporting by The Batch, Analytics DeepLearning.AI — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Z.ai just shipped GLM-5.1, an open-weights model built around a simple but consequential idea: don't stop when the token budget runs out, stop when the job is actually done. The model cycles through planning, execution, and self-evaluation, and if it decides its approach isn't working, it changes course and tries again — sometimes racking up thousands of tool calls over multiple hours in the company's own tests. That's a meaningfully different design goal than most models, which wrap up once they hit a length limit or conclude that more reasoning won't help.

The numbers back up the ambition, mostly. On Artificial Analysis' Intelligence Index, GLM-5.1 topped every open-weights model but still sat behind Gemini 3.1 Pro Preview and GPT-5.4, which tied at the top, and behind Claude Opus 4.6. On Arena's Code leaderboard it landed third, just behind two Claude Opus 4.6 configurations. Where it actually led the field, according to Z.ai's own testing, was SWE-Bench Pro, a benchmark built from real GitHub engineering problems, and CyberGym, a cybersecurity reasoning test — though on the latter, Gemini and GPT-5.4 both declined to run certain tasks for safety reasons, which likely dinged their scores. On pure reasoning and math, the gap widens: GLM-5.1 trailed Gemini on GPQA Diamond and GPT-5.4 on AIME 2026 by clear margins.

Z.ai also raised prices meaningfully. API token costs are up roughly 40 percent over the prior GLM version, and coding-plan subscriptions have roughly doubled. It's still cheaper per million input tokens than Claude Opus 4.6, but that gap is shrinking fast, which says something about where Z.ai thinks this capability sits in the market.

The real test isn't the benchmark table, though — it's whether GLM-5.1's persistence translates into actually finishing hard, messy jobs. METR has tracked the length of tasks AI agents can complete autonomously roughly doubling every seven months, and Cursor recently ran a swarm of agents continuously for a week. But newer benchmarks built specifically to probe sustained performance, like SWE-EVO, still show even the best models completing only about a quarter of long-running coding tasks. Knowing when to abandon a dead-end strategy, rather than just grinding longer, is the harder skill, and it's one current scoring systems barely capture.

My take — AI-written commentary, not fact-checked reporting

Z.ai deserves credit for chasing a genuinely underrated capability — knowing when to quit a bad approach — instead of just chasing another leaderboard number, and the pricing jump suggests even Z.ai knows persistence is worth paying for. But until independent testers, not just the company itself, verify those SWE-Bench Pro and CyberGym wins, treat the open-weights crown with a raised eyebrow; self-reported benchmarks have a way of flattering the model that ran them.

Read more about this at: The Batch

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.