Alibaba’s Qwen3-VL Outperforms Gemini 2.5 Pro in Benchmarks

Alibaba’s Qwen3-VL Outperforms Gemini 2.5 Pro in Benchmarks

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
17. 10. 2025
4 minutes reading · 17 views
Alibaba’s Qwen3-VL Outperforms Gemini 2.5 Pro in Benchmarks

Alibaba Cloud introduces Qwen3-VL, a new generation of open multimodal models that combines text and visual content processing. This model focuses on better understanding of images, videos, and text, with an emphasis on longer contexts and agentic interactions. The flagship Qwen3-VL-235B-A22B model is openly available in Instruct and Thinking versions. The Instruct version achieves results comparable to or better than Gemini 2.5 Pro in key visual perception benchmarks, while the Thinking version excels at multimodal reasoning.

Qwen3-VL-235B-A22B-Instruct benchmark

The Qwen3-VL-235B-A22B-Instruct model excels across ten dimensions, including university-level problems, mathematical reasoning, logic puzzles, general visual question answering, subjective experience, instruction following, multilingual text recognition, chart and document parsing, 2D and 3D object localization, multi-image understanding, spatial perception, video understanding, agentic task execution, and code generation. It outperforms closed models such as Gemini 2.5 Pro and GPT-5 on most of these metrics and sets new records among open multimodal models.

Qwen3-VL-235B-A22B-Instruct benchmark tasks

Qwen3-VL-235B-A22B-Thinking is optimized for STEM (science, technology, engineering, and mathematics) and mathematical reasoning. When handling complex questions, it notices subtle details, breaks down problems, analyzes causes and consequences, and provides logical, evidence-based answers. It achieves strong results in benchmarks such as MathVision, MMMU, and MathVista, in some cases even outperforming Gemini 2.5 Pro.

Qwen3-VL-235B-A22B-Thinking benchmark

Key Capability Improvements

Qwen3-VL can operate computer and mobile interfaces, recognize graphical user interface (GUI) elements, understand button functions, call tools, and complete tasks. It achieves state-of-the-art results in the OS World benchmark, where tool use improves performance in fine-grained perception.

The model generates code from images or videos, for example by converting design mockups into Draw.io, HTML, CSS, or JavaScript formats, enabling “what you see is what you get” visual programming.

All models natively support a context of 256K tokens, expandable to up to 1 million tokens. This means they can process hundreds of pages of technical documents, entire textbooks, or two-hour videos, while remembering details and retrieving them accurately, even down to specific seconds in videos.

Optical character recognition (OCR) now supports 32 languages, up from the original 10, and works more reliably under challenging conditions such as low light, blur, or tilted text. Recognition accuracy has improved for rare characters, ancient scripts, and technical terms, as have long-document understanding and fine-structure reconstruction.

Thanks to early joint pretraining of text and visual modalities, the model achieves text capabilities comparable to the Qwen3-235B-A22B-2507 model. This makes it a text-centric multimodal model for the next generation of vision-language systems.

Architecture and Innovations

The architecture retains its native dynamic-resolution design but introduces updates in three areas. The first is Interleaved-MRoPE, where feature dimensions are interleaved across time, height, and width, ensuring full frequency coverage and improving long-video understanding.

The second innovation is DeepStack technology, which integrates multi-level features from ViT (Vision Transformer), improving the capture of visual details and the accuracy of text-image alignment. Instead of inserting visual tokens into only one layer, they are now inserted into multiple layers of the large language model, enabling more fine-grained visual understanding.

The third improvement is a text-timestamp alignment mechanism that replaces the original T-RoPE. It uses a “timestamps - video frames” input format, enabling more precise alignment of temporal information and visual content. The model natively supports output in seconds or in hours:minutes:seconds format, improving the accuracy of action and event localization in videos.

Benchmark Performance

In the “needle in a haystack” test with ultra-long videos, the model achieved 100% accuracy with a context length of 256K tokens. When expanded to 1M tokens, equivalent to approximately two hours of video, accuracy remained at 99.5%.

In a multilingual text recognition test covering languages other than Chinese and English, the model achieved over 70% accuracy in 32 out of 39 languages, demonstrating strong generalization.
Languages

Qwen3-VL-235B-A22B-Instruct supports image-based reasoning with tool use. Tests involving four fine-grained perception and interaction tasks showed consistent performance improvements, confirming that combining image analysis with tool calling enhances visual capabilities.

Additional Models and Availability

In addition to the flagship model, Qwen3-VL-4B and Qwen3-VL-8B are available in Instruct and Thinking versions, suitable for smaller devices. Qwen3-VL-30B-A3B-Instruct and Qwen3-VL-30B-A3B-Thinking are also available, along with FP8 versions for efficient deployment.

The models are compatible with Transformers and vLLM and support quantization and acceleration technologies such as Flash-Attention 2. Developers have access to fine-tuning code and cookbooks for various capabilities, such as omni recognition, document parsing, 2D and 3D localization, OCR, video understanding, and agentic interactions.

Advertisement

Content created with help from UpTier.

SEO and GEO on autopilot. UpTier’s multi-agent systems write and optimize content for search engines and AI answers.

Discover UpTier ↗

Category:AI
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

OpenAI gives Codex reusable cloud workspaces accessible from any deviceOpenAI gives Codex reusable cloud workspaces accessible from any device
Codex gains reusable cloud development environments, alongside voice controls in its CLI, code reviews in the ChatGPT desktop app and cloud-based security tools.
2 min read
2. 10. 2026
Amazon releases Strands Decider 2B for AI workflow decisionsAmazon releases Strands Decider 2B for AI workflow decisions
Strands Decider 2B selects from predefined options and returns a confidence score. The fully open-source model is available now and small enough to run locally.
2 min read
1. 10. 2026
OpenAI says it disrupted a campaign to extract hidden model reasoningOpenAI says it disrupted a campaign to extract hidden model reasoning
OpenAI reported a coordinated effort to extract protected model reasoning and said it closed an extraction pathway. It attributed the main cluster of activity to individuals associated with Moonshot AI, the developer of Kimi.
3 min read
1. 10. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok