TLDRocket
Sign in

An industrial-scale distillation of models, or subtle benchmaxxing: What developers really think of GLM-5.3

The New Stack Adrian Bridgwater Covered by 2 sources

Z.ai just dropped GLM-5.3, claiming big coding gains from training tweaks alone. Developers aren't so sure - one says it looks a lot like distilled Claude.

Based on reporting by The New Stack, Adrian Bridgwater — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Z.ai put out GLM-5.3 on Friday, and the interesting part isn't the model itself so much as how it got made. Same codebase as GLM-5.2, the company says. No new architecture. Every improvement came purely from post-training: reasoning alignment, supervised fine-tuning, and reinforcement learning from human feedback, stacked on top of what was already there. Z.ai's own description, in an unsigned blog post, is that the team spent the past month scaling that same stack with more environments, more varied tasks, and more compute thrown at training on top of them.

What's changed, according to the company, is the texture of the training environments themselves. They now mirror a much wider slice of actual production engineering work, with task categories built around how real teams operate rather than tidy synthetic problems. Some of these tasks are meant to represent several days of an experienced engineer's workload. One example Z.ai gives: an ML infrastructure task where the model gets the same access an engineer would — compute clusters, storage, internal docs, codebases, experiment logs — and has to hunt down bottlenecks, ship optimizations, run experiments, and produce a measurable speedup without breaking anything.

On the numbers side, Z.ai's own Code Bench puts GLM-5.3 at a 50% improvement over GLM-5.2, and the company argues a private benchmark like this dodges contamination from public test sets, giving a cleaner read on real user experience. It also shows results on public benchmarks including TerminalBench 3.0, DeepSWE, Agents' Last Exam, AutomationBench, Humanity's Last Exam with tools, and OpenAI's GDPVal-AA v2. Under the hood, GLM still leans on the General Language Model training approach it's named for, using autoregressive blank infilling — cloze-style tests that mask chunks of data to build up vocabulary, comprehension, and reasoning. Z.ai says it'll release the GLM-5.2 weights two weeks after launch, once safety evaluation and hardening wrap up.

Not everyone buys the framing. Nishant Soni, co-founder of NonBioS.ai, says no benchmark he's seen actually backs up the claimed long-horizon capability in real tasks, and doubts the training setup Z.ai describes could realistically produce datasets diverse enough to build that kind of skill. He goes further, suggesting the real story might be industrial-scale distillation of Anthropic's models — his team has noticed striking similarity between Kimi and GLM outputs and Claude's, in a way that Gemini and Grok don't show.

Other practitioners are less conspiratorial but still cautious. Sherif Higazy of Megaton likes the idea of training environments that push more realism into the model itself, shifting the burden of adaptation upstream instead of leaving it to teams building around the model afterward — though he admits not everyone's sold on the approach yet. Rohan Kodialam of Sphinx takes the widest view: with OpenAI, Anthropic, Google, and Chinese open-weight labs constantly leapfrogging each other, he argues enterprises should keep business context separate from any one model and avoid irreversible choices, since compute availability alone still tilts the scales toward the bigger US labs over newer entrants like Z.ai.

ML researcher Nathan Lambert offers maybe the sharpest framing of why Chinese labs keep pace at all. He doesn't think there's heavy benchmaxxing going on, just a subtle version of it — and points out that American labs likely have far stronger models sitting unreleased for months while they test internally. That gap, he argues, is exactly the window Chinese labs use to keep climbing the public leaderboards before the next US release resets the comparison.

My take — AI-written commentary, not fact-checked reporting

Post-training gains are cheap to claim and hard to independently verify, which is precisely why an outside skeptic pointing at output similarity to Claude matters more than another in-house benchmark chart. If the real edge here is that US labs sit on stronger models for months before shipping, the fair fight isn't happening at all - it's a timing game, and Chinese labs are simply better at exploiting the gap. Anyone building serious infrastructure around whichever model tops this week's chart deserves what they get when the leaderboard flips again in six months.

Read more about this at: The New Stack

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.