TLDRocket
Sign in

China’s Low-Priced Z.ai Model Is Exposing Costly Coder Habits

IEEE Spectrum Matthew S. Smith Covered by 7 sources

China's Z.ai just dropped GLM 5.2, a coding model that nearly matches Anthropic's best for a fraction of the price. It's open-weights too, and the real bottleneck might be developer habits, not the tech.

Based on reporting by IEEE Spectrum, Matthew S. Smith — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Zain Hasan splits his coding workload between two tiers: a frontier model like Anthropic's Fable for the hard stuff, and something cheaper for everything else. Lately, that cheaper option has been GLM 5.2, released June 16 by Beijing-based Z.ai. It costs $4.40 per million output tokens through the company's API — less than a fifth of what Anthropic charges for Opus 4.8, and a tenth the price of Fable. The model is also open-weights under an MIT license, so anyone with enough GPUs can just download it and run it themselves.

The benchmark numbers are close enough to spook people. GLM 5.2 nearly ties Opus 4.8 on agentic coding tests like FrontierSWE and PostTrainBench, and it scores well on cybersecurity benchmarks too. But Z.ai's own research report, published the same day as the launch, is more modest than the headlines suggest. It only claims outright wins over Opus 4.8 on two easier reasoning benchmarks, and none in coding. On SWE-Marathon, a brutal long-duration coding test, GLM 5.2 finished just 13 percent of tasks — Opus 4.8 doubled that. Opus also beat it comfortably on NL2Repo, DeepSWE, and Tool-Decathlon.

Where GLM 5.2 seems to actually win is stamina. Hasan says earlier open-weights models would lose coherence after five to fifteen exchanges; this one holds a thread for hours. David Nix, a principal engineer at MetaRouter, now sends 10 to 20 percent of his daily LLM work to GLM 5.2, especially front-end tasks, and calls it

My take — AI-written commentary, not fact-checked reporting

I run open models on my own infrastructure whenever I can, so I'm predisposed to like this — but the real story here isn't China catching up, it's that most engineers still don't check the price tag before they hit send. Whoever's paying the bill is the one actually competing with Anthropic and OpenAI, not the model itself. Give teams a real token budget and watch how fast 'good enough and cheap' starts beating 'best in class.'

Read more about this at: IEEE Spectrum

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.