Alibaba Unveils Qwen-Image: AI That Paints with Words and Colors

Alibaba Unveils Qwen-Image: AI That Paints with Words and Colors

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
6. 8. 2025
4 minutes reading
Alibaba Unveils Qwen-Image: AI That Paints with Words and Colors

Alibaba Introduces Qwen-Image: AI That Paints with Words and Colors

Qwen-Image is the latest artificial intelligence model from Alibaba Cloud, specializing in image generation and editing. This 20-billion-parameter model is built on the MMDiT architecture and delivers significant advances in complex text rendering and precise image editing. According to the official announcement prepared by the Qwen Team on August 4, 2025, the model is available on platforms such as GitHub, Hugging Face, and ModelScope. You can try its demo version on Qwen Chat by simply selecting the image generation option. The model is open-source under the Apache 2.0 license, which means anyone can freely use and modify it.

Key features include excellent text rendering, consistent image editing, and strong performance across various benchmarks. For example, the model can handle complex textual elements in images, such as multi-line layouts, paragraphs, and fine details, in both alphabetic and logographic languages. This makes it an ideal tool for content creators who need accurate and realistic results.

Performance and Benchmarks

Qwen-Image was tested on several public benchmarks, where it achieved state-of-the-art results. In image generation tests such as GenEval, DPG, and OneIG-Bench, it outperformed existing models. For image editing, it excels in GEdit, ImgEdit, and GSO. In text rendering specifically, including LongText-Bench, ChineseWord, and TextCraft, the model significantly outperforms the competition, particularly when generating text in logographic languages. These results show that Qwen-Image is a powerful foundation model for a wide range of visual content tasks.

According to evaluations from various sources, such as a technical overview on YouTube, the model requires substantial hardware resources—up to 57 GB of VRAM (graphics card memory) to run large versions locally. This means it is best suited to users with powerful hardware, but thanks to its open-source availability, it can be integrated into various applications. Testers praise its ability to preserve semantic meaning and visual realism during edits, making it suitable for both creative and analytical tasks.

Benchmarks

Examples of Text Rendering

One of Qwen-Image's greatest strengths is its ability to render text in various scenarios. For example, in a Hayao Miyazaki-style anime scene, the model generated an image featuring signs on shops and containers related to cloud services, such as cloud storage, cloud computing, and cloud models. The characters have accurate poses and expressions, and the depth of field is realistic.

Shops

Another example features traditional couplets hanging in a Chinese room with blue-and-white porcelain and a painting of a famous tower. The model accurately applied a calligraphic effect and generated details such as flowers and architecture.

Room

In examples featuring an alphabetic language, the model successfully created a bookstore window display with signs about this week's new releases, bestsellers, and the titles of four books. It even handled a complex infographic slide with six sections on emotional well-being, where each module included an icon, heading, and descriptive text, such as a mindfulness practice with a sentence about being present and observing without judgment.

Bookstore

It even supports bilingual rendering with alternating languages, as in an example describing the model as a powerful foundation model for complex text rendering and precise image editing.

Image Editing and Other Capabilities

Qwen-Image is not only about generation, but also editing. It supports operations such as style transfer, adding or removing elements, enhancing details, editing text, and adjusting character poses. This allows users to achieve professional results without complex software. The model also handles a wide range of artistic styles, from photorealistic scenes to impressionist paintings or anime.

Image Editing

One example is the creation of a movie poster with a title about unleashing the imagination, a subtitle about entering a world beyond imagination, and cast and director credits associated with the model. The central visual features a futuristic computer bursting with colors and creatures, in 32K resolution with ultra-fine details.

User Reviews and Comments

According to available reviews and comments, Qwen-Image is praised for its open-source availability and robust handling of complex prompts. Testers commend its photorealistic outputs, high level of detail, and semantic consistency, particularly in multilingual text, where Western models often fail. For example, a Cybernews review highlights its superiority in benchmarks such as GenEval and DPG, as well as its ability to render text with high fidelity.

Users on platforms such as YouTube note that the model is ideal for creative professionals and researchers thanks to its advanced decoder, which produces complex textures and precise lighting. However, some comments point out limitations, such as high hardware requirements—up to 57 GB of VRAM for large models. Overall, Qwen-Image is seen as a versatile tool that lowers barriers to visual content creation and supports innovation in generative AI.

In summary, Qwen-Image represents a step forward in artificial intelligence for images, with an emphasis on accuracy and creativity. The Qwen Team hopes the model will support the community and enable new applications.

Category:AI
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

Altman Announced the Singularity Days After His Models Escaped the Lab on Their OwnAltman Announced the Singularity Days After His Models Escaped the Lab on Their Own
OpenAI chief Sam Altman declared on the Relentless podcast that humanity has already entered the singularity. “We’re like, in the singularity now,” he said verbatim. For decades, the term belonged more to science-fiction literature
6 min read
28. 7. 2026
AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.
Since April, Australian radio has been playing a dance remake of Madonna’s hit Like a Prayer on repeat. Released by Queensland DJ Josh Fawaz, it tops the radio airplay chart and has 35 million Spotify streams.
6 min read
28. 7. 2026
Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?
A first-person shooter that runs directly in the browser, with its own physics and eleven separate code modules. Around 55,000 lines in total, split across eleven subsystems and built on Thr
4 min read
28. 7. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok