TLDRocket
Sign in

Notes on Qwen-Max-0428

GitHub Pages

Alibaba's Qwen team dropped a new flagship chat model, Qwen-Max-0428. It just cracked the top 10 on Chatbot Arena, beating their own previous best.

Based on reporting by GitHub Pages — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Alibaba's Qwen team has quietly pushed out an upgrade to its biggest chat model, and the numbers back up the bragging rights. Qwen-Max-0428 scored 8.96 on MT-Bench, edging past Qwen1.5-110B-Chat's 8.88, and it climbed to 1186 on the Chatbot Arena leaderboard, good enough for a top-10 spot. Small jumps on paper, sure, but at the frontier of these leaderboards even a few tenths of a point represents real fighting for position against the likes of GPT-4 and Claude.

What's notable here isn't just the score bump. Qwen-Max-0428 is accessible through DashScope, Alibaba's API platform, and the team has made it compatible with the OpenAI API format. That means anyone already building on OpenAI's client library can point their code at Alibaba's servers with barely any changes — swap the base URL, swap the API key, and you're calling qwen-max instead of gpt-4. That's a deliberate move to lower the switching cost for developers, and it's the kind of interoperability play that matters more for adoption than a leaderboard rank ever will.

The model is also live in a Hugging Face Spaces demo for anyone who wants to poke at it without writing code, and it's baked into Alibaba's own web service and app, though that consumer app is restricted to mainland China users. So outside of China, the API and the demo are really the only doors in.

This release follows the broader Qwen1.5 family launch, which spanned models from 0.5 billion all the way to 110 billion parameters. Qwen-Max-0428 sits above all of them, positioned as the flagship for anyone who wants Alibaba's best rather than the open-weight mid-tier options. No word yet on parameter count or whether weights will ever be released — this one looks firmly closed, API-only, at least for now.

My take — AI-written commentary, not fact-checked reporting

An OpenAI-compatible endpoint is the real headline here, not the MT-Bench bump — it's Alibaba making it trivially easy to defect from OpenAI's ecosystem without rewriting a line of application logic. I'll believe the leaderboard gains once independent red-teaming catches up, but the interoperability trick is the smarter long-term play, and I'd bet more labs copy it before they copy the benchmark chase.

Read more about this at: GitHub Pages

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.