TLDRocket
Sign in

GLM 5.2 and the coming AI margin collapse (part 1)

Martin Alderson Covered by 7 sources

A new open-weights model, GLM5.2, is basically matching Opus and GPT5.5 for coding work at a fraction of the price. That's a real threat to the fat margins frontier AI labs charge on inference, not just training.

Based on reporting by Martin Alderson — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Everyone remembers the DeepSeek panic, when markets decided cheap training meant the AI capex boom was over. That was the wrong lesson. Training is a one-time cost you amortize; inference is where the real money is made, and where the real margins hide. Anthropic and OpenAI charging $25 per million tokens looks a lot less like cost recovery and more like a business built on selling very profitable API calls once the training bill is paid off.

GLM5.2 from Z.ai is the model that finally makes that math uncomfortable for the frontier labs. After weeks of daily use, the author found it genuinely hard to distinguish from Opus in real coding sessions. It thinks a lot, which slows it down for interactive work and isn't ideal, and it has no vision support and weak web search, both real gaps against Claude and GPT. But for background agentic tasks like reviewing pull requests, none of that matters much.

What should worry Anthropic and OpenAI more is how frictionless switching has become. Both Z.ai and Fireworks expose OpenAI- and Anthropic-compatible endpoints, so pointing Claude Code or Codex at GLM5.2 is a one-line config change, not a migration project. There's no vendor lock-in here comparable to swapping out Salesforce. Enterprise data privacy concerns around Z.ai's own China-linked hosting are real, but plenty of third-party providers offer proper contracts, and self-hosting is an option too.

Then there's price. GLM5.2 runs around $4.40 per million tokens, less than a fifth of Opus's retail rate and about 15% of GPT5.5's. It burns more tokens per task since it thinks more, so it's not a perfectly clean comparison, but even accounting for that it's likely over 50% cheaper for most real workflows at comparable quality. And costs are still falling: Wafer's benchmarking found GLM5.2 running 2.75x cheaper per token on AMD hardware versus Nvidia Blackwell, with more serving-stack optimization clearly still to come.

None of this kills the frontier labs outright, but it puts real pressure on the fat margins that made their business model work so cleanly. When switching costs approach zero and quality gaps approach zero, pricing power evaporates fast.

My take — AI-written commentary, not fact-checked reporting

I've been saying for a while that open weights would eventually stop being a curiosity and start being a genuine commercial threat, and GLM5.2 is that moment arriving on schedule. The interesting fight now isn't model quality, it's margin compression, and anyone still pricing inference like it's 2024 is going to get undercut hard. Watch AMD-based serving stacks too; cheaper silicon for inference is the quiet story nobody's pricing in yet.

Read more about this at: Martin Alderson

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.