Qwen3.7-Max Challenges Google for Third Place, AI Saves Whales, Fine-Tuning Breaks Copyright Alignment
The Batch ● Covered by 7 sources
Qwen3.7-Max is now China's top AI model, tying for third in speed among frontier LLMs. But it's closed weight — Alibaba's chasing profit, not open access, with its best model.
Based on reporting by The Batch — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Alibaba's latest flagship, Qwen3.7-Max, arrived this week as the company's pick for heavy-duty agentic work — coding, scientific tasks, the kind of long-running jobs that chew through tokens. It ships text-only, handling up to a million tokens of input and spitting out as many as 64,000 tokens at a clip of roughly 208 tokens per second. Alibaba simultaneously dropped a multimodal sibling, Qwen3.7-Plus-Preview, but it's the Max model drawing the attention. And like every top-tier Qwen release since late 2025, the weights stay locked up.
On Artificial Analysis' Intelligence Index, a composite covering ten economically relevant benchmarks, Qwen3.7-Max lands seventh, just behind Gemini 3.1 Pro Preview and ahead of Gemini 3.5 Flash on high reasoning. On the factual-knowledge test AA-Omniscience it ranks sixth, well behind Gemini 3.1 Pro but ahead of Claude Sonnet 4.6. Its hallucination rate, at 23 percent, was the lowest among the frontier models tested — though that partly comes from refusing to answer more than half the prompts thrown at it, a strategy that flatters accuracy numbers without necessarily helping anyone trying to get a straight answer.
Speed is where it holds its own. Qwen3.7-Max tied for third place in output speed with Gemini 3.5 Flash, trailing only OpenAI's GPT-OSS 120B and GPT-OSS 20B. Alibaba also showed off an unverified agentic demo: over 35 hours, the model made 1,158 tool calls and ran 432 kernel evaluations while optimizing an attention kernel on hardware it hadn't trained on, eventually producing code that ran roughly ten times faster than a standard reference version. It's a striking number, but Artificial Analysis hasn't run its own long-horizon agentic benchmark on the model yet, so treat it as a company-supplied highlight reel rather than independently confirmed proof.
The bigger story sits behind the specs. Qwen3.7-Max and its stablemate Qwen3.6-Max-Preview keep their weights closed, while smaller models like Qwen3.6-27B and Qwen3.6-35B-A3B remain freely available. Alibaba has also started charging for Qwen Code, its command-line coding tool, and the shift has come alongside turnover in the team's leadership. Taken together, it reads like a company deciding its top models are worth more as a product than as a contribution to the open ecosystem.
My take — AI-written commentary, not fact-checked reporting
Alibaba keeping its smaller models open while locking down the flagship is the clearest sign yet that open weights were never a permanent commitment for these labs — they were a market-entry strategy, and once a model is good enough to sell, the door closes. Charging for Qwen Code on top of that makes the direction obvious: this is a company optimizing for revenue, not reach. Anyone who assumed China's AI labs would keep racing to out-open each other should probably stop assuming that.
Read more about this at: The Batch