TLDRocket
Sign in

Kimi K3 vs DeepSeek V4 Pro vs GLM-5.2: Open Trillion-Scale MoE Models Compared on Benchmarks, License, and Serving Cost

MarkTechPost Michal Sutter Covered by 7 sources

Three Chinese labs just dropped trillion-parameter open MoE models: Kimi K3, DeepSeek V4 Pro, GLM-5.2. K3 wins on smarts, DeepSeek wins on price, and only two of them actually let you download the weights yet.

Based on reporting by MarkTechPost, Michal Sutter — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

China's open-weight race just got a third serious entrant, and the numbers are getting silly. Moonshot AI's Kimi K3 clocks in at 2.8 trillion parameters, activating a mere 16 of 896 experts per token, and Moonshot is already calling it the first open 3-trillion-class model. DeepSeek V4 Pro sits at 1.6 trillion total with 49 billion active. Zhipu AI's GLM-5.2, at 744 billion parameters, looks almost modest by comparison — but it held the open-weight crown right up until K3 showed up on July 16.

On Artificial Analysis's neutral Intelligence Index, K3 scores 57 and lands at number three overall, trailing only Claude Fable 5 and GPT-5.6 Sol and sitting in the same tier as Opus 4.8. GLM-5.2 posts 51, DeepSeek V4 Pro trails at 44 on that particular scale, though it flips the script on raw coding: DeepSeek-V4-Pro-Max hits 80.6% on SWE-bench Verified, tying Gemini 3.1 Pro and topping every open-weight model at release. Moonshot's own matched-harness comparisons show K3 beating GLM-5.2 on every shared coding benchmark, sometimes by more than 3x, as in SWE Marathon's 42.0 versus 13.0.

Licensing is where the story gets less flattering for Moonshot. Both DeepSeek V4 Pro and GLM-5.2 are MIT-licensed with full weights already sitting on Hugging Face, ready for anyone to fine-tune or self-host today. K3 is the odd one out: Moonshot has promised weights by July 27 under a Modified MIT license with a light attribution clause that only kicks in past 100 million monthly active users. Until then, K3 is API-only, which matters a lot if your team's whole pitch is

My take — AI-written commentary, not fact-checked reporting

deploying open weights on your own hardware. Cost separates these three about as sharply as the license terms do. At list pricing, a dollar buys roughly 1.15 million output tokens from DeepSeek V4 Pro, about 227,000 from GLM-5.2, and only 67,000 from K3. Artificial Analysis's blended cost-per-task numbers tell the same story: four cents for DeepSeek, thirty-two cents for GLM-5.2, ninety-four cents for K3. GLM-5.2 also wins on raw speed, running near 168 tokens per second against roughly 62 for the other two. Self-hosting any of them is its own headache — GLM-5.2 wants over a terabyte of VRAM, and Moonshot recommends 64-plus accelerators for K3, which puts it firmly in hyperscaler territory rather than homelab territory. So the choice isn't really about which model is "best." It's about which constraint you're optimizing against. Teams chasing peak benchmark numbers and willing to pay API premiums or wait two weeks should look at K3. Teams that care about verifiable, downloadable weights and rock-bottom serving costs today have a clear answer in DeepSeek V4 Pro, with GLM-5.2 as the fast, self-hostable middle ground that punches above its parameter count.","opinion":"I'll say the quiet part: a model that scores well but isn't actually downloadable yet doesn't count as \"open\" in any way that matters to a procurement team, it counts as a really good ad. DeepSeek shipping MIT weights on day one while undercutting everyone on price is the more consequential story here, benchmark leaderboard be damned. And watch that July 27 date — if Moonshot slips it, K3's whole \"first open 3T model\" claim starts looking like vaporware with a really nice API."}

Read more about this at: MarkTechPost

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.