TLDRocket
Sign in

GLM-5.2 Is The New Best Open Model

TLDR Dev Covered by 3 sources

GLM-5.2 just dropped and it's the strongest open-weight model yet, trading blows with Opus 4.7 on many benchmarks. It's also almost certainly distilled from Claude, which explains both the shine and the catch.

Zhipu's GLM-5.2 landed last week and immediately started climbing leaderboards that open models rarely touch. Artificial Analysis puts it at 51 on their combined index, trailing only Fable, Opus 4.8, GPT-5.5 and Opus 4.7, and tied with GPT-5.4. On FrontierSWE it sits one notch behind Opus 4.8 and one ahead of GPT-5.5. On PosttrainBench it actually takes the top spot, edging out Opus 4.8. That's a serious jump from GLM-5.1, and by most measures a bigger leap toward the frontier than DeepSeek R1 managed during its own viral moment last year.

But the benchmarks flatter it more than real use does. Dig into evaluations that are harder to game — Arena text rankings, the agent leaderboard, sycophancy tests — and GLM-5.2 slides toward the middle of the pack rather than the top. That pattern lines up with a fairly obvious tell: the model frequently identifies itself as Claude, speaks in Claude's cadence, and runs happily inside a Claude-style harness. The working theory among people who've poked at it, including longtime China-AI watcher Teortaxes, is that it's heavily distilled from Claude Opus. Distillation isn't cheating exactly, but it does mean the model shines on tasks that resemble its training data and gets shakier the moment you wander off that path.

Users testing it in the wild are genuinely split, and both camps have receipts. Jeremy Howard called it a marvel that handles long context better than anything he's tried in open weights. A Kubernetes engineer had it debug an Envoy Gateway issue that stumped Opus 4.8, for $7.32 in API spend, complete with the model noticing it was redoing a fix Claude had previously reverted. On the other side, Theo from t3.gg points out that Opus 4.8 and GPT-5.5 on medium settings are both cheaper and smarter, and that GLM-5.2 burns through far more output tokens per answer, which means longer waits even at a lower per-token price. There's also no native vision support, which multiple testers flagged as a genuine gap for anything involving images or dashboards.

That leaves GLM-5.2 in an oddly narrow lane. It's not cheap enough to beat lighter open models on bulk, high-volume work, and it's not strong enough to beat the closed frontier on the hardest problems. Where it wins is a specific middle ground: strong coding and agentic tasks at a lower cost than Opus or GPT, for people who either need open weights specifically or are willing to trade some reliability and token efficiency for that openness. Best estimates put it something like four to seven months behind the actual frontier, ignoring the missing vision and the generalization concerns baked in from distillation — which is close, but not a repeat of the DeepSeek moment so much as the newest entry in the ongoing 'closest open model yet' cycle.

My take

I'll say what a lot of people are dancing around: a model that thinks it's Claude and behaves like Claude in a harness built for Claude isn't really proof China closed the gap — it's proof someone found a very good photocopier. That's still useful, and cheaper agentic coding is genuinely good news for builders who can't afford frontier API bills. But treating this as the new DeepSeek moment undersells how much of the gap between open and closed AI is now measured in distillation lag rather than raw research progress, and that's a much less exciting story than the headlines suggest.

Read more about this at: TLDR Dev

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.