Qwen-Image: Crafting with Native Text Rendering
Qwen
Alibaba released Qwen-Image, a 20 billion parameter multimodal diffusion transformer model that specializes in text rendering and image editing across multiple languages and scripts. The model achieved state-of-the-art performance on 9 public benchmarks including GenEval, DPG, OneIG-Bench, GEdit, ImgEdit, GSO, LongText-Bench, ChineseWord, and TextCraft. The model enables users to generate images with precise text, edit existing images while preserving semantic meaning, and create professional content like posters and presentations with complex multilingual layouts.
Why it matters
GITHUB HUGGING FACE MODELSCOPE DEMO DISCORD We are thrilled to release Qwen-Image, a 20B MMDiT image foundation model that achieves significant advances in complex text rendering and precise image editing. To try the latest model, feel free to visit Qwen Chat and choose “Image Generation”. The key features include: Superior Text Rendering: Qwen-Image excels at complex text rendering, including multi-line layouts, paragraph-level semantics, and fine-grained details. It supports both alphabetic languages (e.