Qwen3.8-Flash-Next
Simon Willison’s Weblog Simon Willison ● Covered by 4 sources
Qwen released Qwen3.8-Flash-Next, an open-weights multimodal Mixture-of-Experts model. It is trained on 125B tokens with 6B active parameters. The model is positioned as an early preview of the architecture intended for Qwen4, and users report running it on a DGX Spark with quantized variants up to about 78.9GB.
Why it matters
Qwen3.8-Flash-Next Another open weights model from Qwen. This one is "a multimodal MoE model that also serves as an early preview of the architecture used in Qwen4". It's pretty big: 125B tokens, but only 6B active which means it gets a significant performance boost. I've been trying it out on a DGX Spark using these Unsloth quantized models. I'm still exploring the model - so far I've tried the 72.5GB UD-IQ1_S one (producing these pelicans) and the 78.9GB UD-Q2_K_XL (producing these). My favorite so far was this xhigh reasoning effort one from UD-Q2_K_XL: Via Hacker News Tags: ai, generative-ai, llms, qwen, pelican-riding-a-bicycle, ai-in-china, nvidia-spark