TLDRocket
Sign in

[AINews] Qwen 3.8 Max(2.4T) and 27B, new open weights models for Coding and Cowork

Latent Space Covered by 6 sources

Alibaba just dropped Qwen3.8-Max, a 2.4 trillion parameter open model that's already beating Claude and GPT-5.5 on coding benchmarks. Open weights land next week, but running this thing yourself needs serious hardware.

Alibaba's Qwen team just answered the question everyone's been asking since last year's management shakeup: is this lab still serious about open weights? The answer is a 2.4 trillion parameter model called Qwen3.8-Max, and it's not messing around. On Vals AI's Index it scored 66.1, tying Claude Opus 4.7 exactly, while costing about 2.3 times less per test. On SWE-bench it hit 87.3%, ahead of GPT-5.5's 82.6% and GLM-5.2's 83.3%, trailing only Claude Opus 4.8 at 89.2%. Open weights for both this flagship and a smaller 27B sibling are promised for next week.

The headline demos read like science fiction with receipts. Alibaba says the model ran unattended for more than 10 days straight while building its own coding harness, then spent 125 hours autonomously rebuilding and improving a published data-selection research paper, beating the original benchmark by 2.71 points. It also entered a real competition against 526 human teams at WWW2025 and landed in the top 13% within 24 hours. There's a chip design flow in there too, where it took a cryptographic accelerator layout from 8,298 gates down to 678, shrinking die area by 81% while hitting 500MHz timing closure. Whether every claim survives independent scrutiny is another matter, but the sheer breadth is a statement of intent.

Here's the catch nobody should skip past: open weights doesn't mean anyone can actually run this on a laptop. Jamin Ball pointed out that Qwen3.8-Max, like Kimi K3's 104B active parameters or GLM-5.2's 744B total, demands serious infrastructure — Moonshot recommends 64-plus accelerators for K3-class models. Qwen3.8-Max reportedly activates only about 95B of its 2.4T parameters per token, a roughly 4% activation ratio that explains how Alibaba can price API access at $2 input and $6 output per million tokens. But sparse MoE efficiency at inference doesn't erase the fact that loading the weights alone requires enterprise-grade GPU clusters, not a gaming rig.

Then there's the license question, which got messier than the benchmarks. OstrisAI flagged terms that appeared to restrict downloading the model from the US, EU, UK, and Korea — echoing similar complaints aimed at MiniMax's H3 release. Alibaba hasn't issued a clarifying statement in the threads collected here, so the restrictive reading stands unresolved. For engineers,

My take

,

Read more about this at: Latent Space

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.