TLDRocket
Sign in

GLM-5.3 didn’t change the base model — where did its coding gains come from?

The New Stack Amanda Caswell Covered by 3 sources

Z.ai’s GLM-5.3 keeps the same base as GLM-5.2, but its coding scores jumped. The boost came mostly from heavier post-training, not a new foundation model.

Based on reporting by The New Stack, Amanda Caswell — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Z.ai shipped GLM-5.3 on Friday, and the odd part is what didn’t change: the base model is the same one used for GLM-5.2. The gains came later, in post-training, where Z.ai says it pushed the model through far more long-horizon task environments and more developer tooling and engineering workflows than before.

That training setup sounds less like textbook model tuning and more like a rehearsal for an actual software team. Some of the tasks covered the whole cycle: find the bug, write the fix, run the tests, ship the result. Z.ai says a single task could resemble the workload of a senior engineer spread over several days. It’s a pretty blunt signal that the company is betting on compute spent in the right places, not just bigger parameter counts.

The numbers back up the story, at least for now. Z.ai says GLM-5.3 is up 50% on its internal Code Bench versus GLM-5.2. On public tests, the model moved from 4.6 to 28.3 on Terminal-Bench 3.0, from 46.2 to 66.9 on DeepSWE v1.1, and from 23.8 to 28.5 on Agents’ Last Exam. That DeepSWE result sits close to Google’s Gemini 3.7 Flash at 65%, though different test harnesses make tidy one-to-one comparisons slippery.

Developers can already use GLM-5.3 through Z.ai’s Coding Plan with Claude Code, Cline, OpenCode and Codex, but direct API access is still marked as coming soon. Z.ai says it plans to release the weights after two weeks of hardening and safety testing. Until then, teams get a model with a 1-million-token context window, a 128,000-token completion ceiling and configurable reasoning effort levels, but not yet the cleanest setup for local testing or apples-to-apples comparisons.

Security is part of the pitch too. Z.ai says GLM-5.3 scored 84.5% on CyberGym, up from 77.2% for GLM-5.2, but its weaker showing on ExploitBench suggests it is better at finding and reviewing flaws than turning them into working attacks. That split matters. A model that can read code well is useful; one that can reliably exploit production systems is a different beast entirely.

My take — AI-written commentary, not fact-checked reporting

This is the part of AI progress that actually matters: not a shinier base model, but better post-training where the work gets done. The industry loves to brag about bigger models because it sounds expensive and impressive; the cheaper trick is to make the model useful in the mess where developers live. That’s the version worth watching, and also the one most likely to get hidden behind a “coming soon” button.

Read more about this at: The New Stack

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.