Alibaba’s Qwen3-VL Outperforms Gemini 2.5 Pro in Benchmarks

Alibaba’s Qwen3-VL Outperforms Gemini 2.5 Pro in Benchmarks

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
17. 10. 2025
4 minutes reading
Alibaba’s Qwen3-VL Outperforms Gemini 2.5 Pro in Benchmarks

Alibaba Cloud introduces Qwen3-VL, a new generation of open multimodal models that combines text and visual content processing. This model focuses on better understanding of images, videos, and text, with an emphasis on longer contexts and agentic interactions. The flagship Qwen3-VL-235B-A22B model is openly available in Instruct and Thinking versions. The Instruct version achieves results comparable to or better than Gemini 2.5 Pro in key visual perception benchmarks, while the Thinking version excels at multimodal reasoning.

Qwen3-VL-235B-A22B-Instruct benchmark

The Qwen3-VL-235B-A22B-Instruct model excels across ten dimensions, including university-level problems, mathematical reasoning, logic puzzles, general visual question answering, subjective experience, instruction following, multilingual text recognition, chart and document parsing, 2D and 3D object localization, multi-image understanding, spatial perception, video understanding, agentic task execution, and code generation. It outperforms closed models such as Gemini 2.5 Pro and GPT-5 on most of these metrics and sets new records among open multimodal models.

Qwen3-VL-235B-A22B-Instruct benchmark tasks

Qwen3-VL-235B-A22B-Thinking is optimized for STEM (science, technology, engineering, and mathematics) and mathematical reasoning. When handling complex questions, it notices subtle details, breaks down problems, analyzes causes and consequences, and provides logical, evidence-based answers. It achieves strong results in benchmarks such as MathVision, MMMU, and MathVista, in some cases even outperforming Gemini 2.5 Pro.

Qwen3-VL-235B-A22B-Thinking benchmark

Key Capability Improvements

Qwen3-VL can operate computer and mobile interfaces, recognize graphical user interface (GUI) elements, understand button functions, call tools, and complete tasks. It achieves state-of-the-art results in the OS World benchmark, where tool use improves performance in fine-grained perception.

The model generates code from images or videos, for example by converting design mockups into Draw.io, HTML, CSS, or JavaScript formats, enabling “what you see is what you get” visual programming.

All models natively support a context of 256K tokens, expandable to up to 1 million tokens. This means they can process hundreds of pages of technical documents, entire textbooks, or two-hour videos, while remembering details and retrieving them accurately, even down to specific seconds in videos.

Optical character recognition (OCR) now supports 32 languages, up from the original 10, and works more reliably under challenging conditions such as low light, blur, or tilted text. Recognition accuracy has improved for rare characters, ancient scripts, and technical terms, as have long-document understanding and fine-structure reconstruction.

Thanks to early joint pretraining of text and visual modalities, the model achieves text capabilities comparable to the Qwen3-235B-A22B-2507 model. This makes it a text-centric multimodal model for the next generation of vision-language systems.

Architecture and Innovations

The architecture retains its native dynamic-resolution design but introduces updates in three areas. The first is Interleaved-MRoPE, where feature dimensions are interleaved across time, height, and width, ensuring full frequency coverage and improving long-video understanding.

The second innovation is DeepStack technology, which integrates multi-level features from ViT (Vision Transformer), improving the capture of visual details and the accuracy of text-image alignment. Instead of inserting visual tokens into only one layer, they are now inserted into multiple layers of the large language model, enabling more fine-grained visual understanding.

The third improvement is a text-timestamp alignment mechanism that replaces the original T-RoPE. It uses a “timestamps - video frames” input format, enabling more precise alignment of temporal information and visual content. The model natively supports output in seconds or in hours:minutes:seconds format, improving the accuracy of action and event localization in videos.

Benchmark Performance

In the “needle in a haystack” test with ultra-long videos, the model achieved 100% accuracy with a context length of 256K tokens. When expanded to 1M tokens, equivalent to approximately two hours of video, accuracy remained at 99.5%.

In a multilingual text recognition test covering languages other than Chinese and English, the model achieved over 70% accuracy in 32 out of 39 languages, demonstrating strong generalization.
Languages

Qwen3-VL-235B-A22B-Instruct supports image-based reasoning with tool use. Tests involving four fine-grained perception and interaction tasks showed consistent performance improvements, confirming that combining image analysis with tool calling enhances visual capabilities.

Additional Models and Availability

In addition to the flagship model, Qwen3-VL-4B and Qwen3-VL-8B are available in Instruct and Thinking versions, suitable for smaller devices. Qwen3-VL-30B-A3B-Instruct and Qwen3-VL-30B-A3B-Thinking are also available, along with FP8 versions for efficient deployment.

The models are compatible with Transformers and vLLM and support quantization and acceleration technologies such as Flash-Attention 2. Developers have access to fine-tuning code and cookbooks for various capabilities, such as omni recognition, document parsing, 2D and 3D localization, OCR, video understanding, and agentic interactions.

Category:AI
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

Altman Announced the Singularity Days After His Models Escaped the Lab on Their OwnAltman Announced the Singularity Days After His Models Escaped the Lab on Their Own
OpenAI chief Sam Altman declared on the Relentless podcast that humanity has already entered the singularity. “We’re like, in the singularity now,” he said verbatim. For decades, the term belonged more to science-fiction literature
6 min read
28. 7. 2026
AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.
Since April, Australian radio has been playing a dance remake of Madonna’s hit Like a Prayer on repeat. Released by Queensland DJ Josh Fawaz, it tops the radio airplay chart and has 35 million Spotify streams.
6 min read
28. 7. 2026
Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?
A first-person shooter that runs directly in the browser, with its own physics and eleven separate code modules. Around 55,000 lines in total, split across eleven subsystems and built on Thr
4 min read
28. 7. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok