Alibaba’s Qwen team published Qwen-Image-2.1 on Hugging Face over the weekend of 19 and 20 September: a combined text-to-image and image-editing model whose visual generation component has 7 billion parameters across 32 single-stream DiT layers.
The model card sets out four changes. A lighter architecture using mixed-granularity attention and prefix KV cache reuse. Native transparency, so the model generates and edits RGBA images and can lift subjects out of photographs. Editing that accepts up to ten reference images, with local edits marked by circles or masks. And improvements to typography, portrait lighting and fine detail. Output runs to 2752 by 1536 pixels.
The detail worth pausing on is the licence. These weights are released under the Qwen Research License Agreement, where the repository’s earlier Qwen-Image releases were Apache 2.0. Open weights and open source are not the same thing, and a research licence is precisely the gap between them: you can download the model, but the terms are the vendor’s rather than the OSI’s.
The card describes capabilities rather than publishing comparative scores, so for now anyone ranking 2.1 against closed systems is working from their own testing rather than from Qwen’s.
