TLDRocket
Sign in

Qwen3.8-Max is The Next Chinese Open-Weights Assault on The AI Frontier

Trending Topics Jakob Steinschaden Covered by 10 sources

Alibaba just launched Qwen3.8-Max, a massive new AI model, and it's about to open-source its biggest one yet. It crushes rivals on images and documents but still lags on hard coding tasks.

Based on reporting by Trending Topics, Jakob Steinschaden — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Alibaba dropped its biggest AI model yet on Monday, and this one comes with a twist nobody was expecting: the company plans to release the weights next week, making it the first time a model from its top-tier Max line has gone open. Until now, Alibaba kept its strongest work locked behind an API and only handed out the mid-tier stuff for free. Qwen3.8-Max, with 2.4 trillion total parameters and 95 billion active per step thanks to its mixture-of-experts design, breaks that pattern. A smaller sibling, Qwen3.8-27B, is also getting freely released.

The headline numbers are wild if you take them at face value. Alibaba says the model spent more than ten straight days building a software project called oh-my-cli, cranking out 265 commits and 127 pull requests entirely on its own. In another test, it ran a simulated year of trading on Taobao and Tmall data and turned 100,000 yuan into 416,252 yuan. It also redesigned a cryptographic chip, slashing logic gates from 8,298 down to 678 across roughly 500 steps. These are the kind of long-horizon, autonomous-agent tasks Alibaba is clearly betting on, and it even trained the model to run inside rival agent frameworks like Anthropic's Claude Code and OpenAI's Codex, not just its own QwenWork.

But coding is where the story gets less flattering. Across twelve coding benchmarks, Qwen3.8-Max only comes out on top in one, PaperBench, while Claude Fable 5 and GPT-5.6 Sol take most of the rest. Where the model actually shines is everything visual: image, video and document processing. On a visual perception test called Dense200, it scored 87.0 versus Gemini 3.1 Pro's 69.7 and Claude Fable 5's mere 31.1. It also topped video understanding and swept every single document and office task in the comparison, roughly two-thirds of the 54 tests in that category overall.

A few caveats make the numbers harder to trust blindly. Claude Opus 5, currently the top scorer on Artificial Analysis's Intelligence Index, is missing entirely from Alibaba's comparison table, replaced by the older Opus 4.8. Several benchmarks are Alibaba's own creations with no outside verification, and some results for Claude Fable 5 may reflect Anthropic's automatic model-swapping on certain topics rather than Fable 5 itself. There's no independent listing yet on Artificial Analysis or OpenRouter, and no model card detailing safety testing or error rates.

On pricing, Qwen3.8-Max lands at 2 dollars per million input tokens and 6 dollars per million output tokens, a jump from its predecessor's discounted preview rate but still far cheaper than Claude Fable 5's 10-and-50 or GPT-5.6 Sol's 5-and-30. Since the model appears to think at length by default, generating extra billable output tokens along the way, nobody yet has a real cost-per-task figure to compare against competitors.

My take — AI-written commentary, not fact-checked reporting

Another Chinese lab is about to hand out weights for a frontier-class model while Anthropic and OpenAI keep their best work locked in a paywalled API, and that gap is the real story here, not the cherry-picked benchmarks. Self-graded scorecards with missing competitors and homemade tests are marketing, not science, so wait for Artificial Analysis to run its own numbers before believing the hype. Still, if Alibaba actually ships those weights next week, it's one more data point that open-weights momentum keeps building in Beijing while the West debates safety papers.

Read more about this at: Trending Topics

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.