TLDRocket
Sign in

Z.ai Ships GLM-5.3 Without Retraining the Base Model: Better at Complex Coding and Long-Horizon Tasks

MarkTechPost Asif Razzaq Covered by 3 sources

Z.ai shipped GLM-5.3 without retraining the base model. It’s stronger on long coding runs, and the security gains were bigger than expected.

Based on reporting by MarkTechPost, Asif Razzaq — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Z.ai has released GLM-5.3, and the interesting part is what it did not change. The model runs on the same 743B base as GLM-5.2. The lift came from post-training: more task environments, more environment types, and more time spent training on them.

That shows up most clearly in coding. On Terminal-Bench 3.0, the score jumps from 4.6 to 28.3. DeepSWE v1.1 rises from 46.2 to 66.9, and Agents’ Last Exam (CLI) moves from 23.8 to 28.5. Z.ai also says its internal Z.ai Code Bench shows a 50% improvement over GLM-5.2, with 31.4% at roughly 50,000 output tokens per task. By comparison, it cites Claude Opus 4.8 at 29.5% with 120,000 tokens, while Claude Fable 5 remains ahead at 39.5% at maximum effort.

The company’s cybersecurity results are even more eye-catching because Z.ai says they were not the main target. It added vulnerability-discovery data expecting better single-bug reasoning, but instead saw the model start building coherent plans across full exploitation chains. On CyberGym, which covers discovery and validation from white-box source, GLM-5.3 moves from 77.2% to 84.5%, slightly ahead of Mythos 5 at 83.8% and GPT-5.6 Sol at 83.6%.

ExploitBench tells a rougher story and a more useful one. GLM-5.3 climbs from 24.4% to 54.4% on a benchmark that needs root-cause reasoning and a working exploit, though Mythos 5 is still well out front at 78.0%. On ExploitGym, GLM-5.3 completes 105 tasks in two hours and 130 in six, versus 29 and 39 for GLM-5.2. The pattern Z.ai is pointing to is simple: the deeper the benchmark goes into the exploitation chain, the bigger the jump.

The model is already live through the Z.ai API, the GLM Coding Plan, and ZCode, but the weights are still private. Z.ai says they should arrive about two weeks after launch, once safety evaluation and hardening are done. That makes this a real release for teams that can use hosted access now, and a wait-and-see for anyone who needs their own weights before touching production.

My take — AI-written commentary, not fact-checked reporting

This is the part of AI that actually matters: not another bigger base model, but a company squeezing real gains out of training discipline. The awkward bit is that the sharpest improvements are showing up in long-horizon coding and exploit chains, which is exactly where people will be tempted to move fastest. Open weights after safety checks sound tidy enough, but the industry keeps proving it likes shipping capability first and thinking about the blast radius right after.

Read more about this at: MarkTechPost

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.