TLDRocket
Sign in

Alibaba’s new model promises Opus 4.6-level performance on your laptop

The New Stack Frederic Lardinois

Alibaba opened a 27B Qwen3.8 model you can run on a laptop. Its benchmarks put it close to Anthropic’s Opus 4.6, with vision and video too.

Based on reporting by The New Stack, Frederic Lardinois — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Alibaba has put the open weights of Qwen3.8 out in the world, and the headline number is absurd: 2.4 trillion parameters. That version is aimed squarely at the frontier, with benchmarks that line it up against closed models from American labs. But the more practical release landed on Friday: a dense 27 billion parameter Qwen3.8 model under Apache 2.0.

That smaller model is the one people can actually imagine using. On a well-specced MacBook Pro or Mac Studio, it should run locally, and Alibaba’s own numbers say it sits in the same performance band as Anthropic’s Opus 4.6 at Max. On several tests it goes further, especially on computer use, coding, and knowledge work. It also handles vision, including video, which gives local users something more interesting than yet another text box with delusions of grandeur.

Alibaba says Qwen3.8-27B beats the prior Qwen 3.7-Plus, with the biggest jump showing up in coding and knowledge work. One example stands out: DeepSWE agentic coding scores move from 14.2 to 42.2. That still trails the current frontier, and Alibaba also says the model lags newer open releases like GLM-5.3. But it is close enough to matter, especially when the comparison set includes models that are not going onto your laptop any time soon.

There are caveats. Benchmarks do not always survive contact with the real world, and for agentic tasks the surrounding harness can matter almost as much as the model. Early reports also say the model can overthink. And the numbers Alibaba publishes are for the original checkpoint, not the quantized versions most local users will actually run, where quality usually gives up something to gain practicality.

The hardware math is the real bottleneck. The unquantized repository weighs 55.6GB before runtime and cache. Community MLX conversions for Apple silicon are already out, though: about 16.1GB for 4-bit and 29.5GB for 8-bit. Alibaba says the model supports 262,000 tokens of context by default, and it plans to extend its hosted version to 1 million tokens with YaRN. On a desktop, that headline number and the usable number are not the same thing.

My take — AI-written commentary, not fact-checked reporting

The open-model crowd keeps proving the same point: if you hand people enough parameters and a decent license, they’ll do the rest on their own machines. The funny part is that the real competition now isn’t just closed labs — it’s memory, heat, and whether your laptop fan sounds like a small aircraft.

Read more about this at: The New Stack

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.