Alibaba has released Qwen-Image-2.1, combining text-to-image generation and instruction-based editing in a single downloadable checkpoint. It also supports transparent output: compatible pipelines can save PNGs with an alpha channel directly, without a separate background-removal step. The model’s image backbone has 7 billion parameters.
Native 2,048 × 2,048 resolution and up to ten reference images
According to Alibaba’s announcement, the model generates images at a native resolution of 2,048 × 2,048 pixels. It accepts up to 10 reference images per generation, allowing input to extend beyond a text prompt or a single visual reference.
RGBA support provides transparency. Alongside the red, green and blue components, the autoencoder retains an alpha channel. Saving transparent PNGs directly depends on whether the chosen pipeline supports that capability.
Three components and an approximately 33 GB download
The image backbone uses a single-stream diffusion transformer with 32 layers and 7 billion parameters. Qwen3-VL 8B serves as the text encoder. The image codec is an RGBA variational autoencoder with 64 channels and 16× spatial compression.
The complete package, covering the image model, encoder and autoencoder, is approximately 33 GB. Runtime memory requirements depend on precision, resolution and quantization. Other factors include component offloading and whether the pipeline keeps all components loaded in GPU memory at once.
First among open-weight models, 18th overall on Artificial Analysis
Qwen-Image-2.1 ranks first among open-weight models on Artificial Analysis’s AA-Image-T2I v2.0 and AA-Image-Editing v2.0 leaderboards. In text-to-image generation, Artificial Analysis gives it an Elo rating of 1,034, compared with 1,011 for Ideogram 4.0 (Quality). For editing, its rating is 1,074, against 1,066 for HunyuanImage 3.0 Instruct.
These Elo ratings turn pairwise preference comparisons into relative scores. The numbers are comparable only within the same leaderboard and can change as further evaluations are added.
Including proprietary systems, Artificial Analysis places the new model 18th in both disciplines. The previous Qwen Image 2.0 ranked 72nd in generation and 58th in editing. Closed systems still dominate the overall standings; on LMArena, Qwen-Image-2.1 ranks 16th in editing and 17th in text-to-image generation.
Measured gains in layout, lighting and text editing
In Artificial Analysis’s category evaluations, the model leads the open-weight field in 16 of 36 categories. It leads in five generation capabilities: complex compositions and text rendering, alongside Knowledge, Layout and Lighting. It leads three generation use-case categories covering social media content, productivity and knowledge work, and consumer use. It ties for the best result in Architecture and Frontier.
For editing, it leads in three types of changes: text or symbol edits, reasoning-based edits, and composition or framing. It also leads five editing use-case categories. These cover productivity, social media content, retail and e-commerce, animation and gaming, and UI/UX design.
Compared with Qwen Image 2.0, Artificial Analysis’s measurements show a smaller gap to the best results across every text-to-image capability. The largest gains are in Layout, Lighting and Knowledge. For editing, the biggest improvements between versions are in Text or Symbol Edits, Object-Level Edit and Identity-Preserving Edit.
In editing, the new version comes closest to the overall leader in Artificial Analysis’s Scene and Style Edit category. That category covers changes to lighting and visual style, as well as background replacement.
Downloadable weights, with a separate license for commercial use
The weights are available through Hugging Face and ModelScope. Diffusers added support on release day, while ComfyUI offers a ready-made workflow template. Support is also available in vLLM-Omni, SGLang and LightX2V.
The downloadable weights come with changed licensing terms. Qwen-Image-2.1 uses the Qwen Research License Agreement, which permits only non-commercial use; commercial deployment requires a separate license from Alibaba. Earlier Qwen-Image releases used Apache 2.0, which allowed commercial use under its terms.



