Qwen3.8-27B: What Alibaba’s New Open Model Can Do and Where It Falls Short

Qwen3.8-27B: What Alibaba’s New Open Model Can Do and Where It Falls Short

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
18. 8. 2026
5 minutes reading · 1 views
Listen to the article
Audio version of the article
Qwen3.8-27B: What Alibaba’s New Open Model Can Do and Where It Falls Short

Alibaba has released Qwen3.8-27B, an open-weight model capable of handling text, images, and video, under the Apache 2.0 license. Its default context window is 262 thousand tokens, and the company claims it can be extended to as many as one million. The release includes its own benchmark table, in which Qwen3.8-27B achieves 61.7 percent on SWE-bench Pro, compared with 53.4 percent for Claude Opus 4.6 Max. Its larger sibling, Qwen3.8-Max, whose weights the team published two days earlier, remains text-only, so the smaller model has attracted the most attention.

What Qwen3.8-27B can do

Qwen3.8-27B is not a text model to which someone has added image processing as an afterthought. The published weights accept text, images, and video as input and return text. The Qwen team also provides runnable image and video examples designed to work with a local copy of the model. This is worth mentioning because the first wave of articles about the Qwen3.8 family described the open weights as text-only. That was true of the Max version from August 12, but not of the 27B model. The configuration file contains a complete vision encoder, and the flag that would be enabled for a purely language model is set to False.

Alibaba’s model targets autonomous assistants running locally. In the announcement, the Qwen team writes that the 27B model outperforms Qwen3.7-Plus overall and excels particularly at real-world programming and office tasks. In practice, this means that a conventional dense model with 27 billion parameters, which fits on a single high-end consumer or workstation graphics card, can now handle long chains of tool calls. Developers previously had to send such requests to the cloud. Alongside the base version, there is also Qwen3.8-27B-FP8 with fine-grained quantization and a block size of 128, which Alibaba says delivers performance nearly identical to the original.

The developers of serving libraries have taken care of speed optimization. The team behind SGLang implemented support on the day of release and reports a decoding speed of 206.1 tokens per second on a single RTX 5090 using the NVFP4 format. On a powerful DGX Spark system, it achieves a throughput of 38.28 tokens per second. The vLLM project also reports rapid integration. 

How it performs in benchmarks

The results table comes from Alibaba, but the company has at least disclosed the testing conditions. It ran all models in the Claude Code environment at a temperature of 1.0, a top_p value of 0.95, and a context length of 256 thousand tokens. The only exception is Claude Opus 4.6 Max on SWE-bench Pro, for which Alibaba used the officially published score.

A benchmark chart titled SWE-Pro, TerminalBench-2.1, and PaperBench compares Qwen3.8-Max with other models.
Benchmark results.

Alibaba claims its widest leads on the QwenSWEBench tests, where it reports 79 percent compared with 63.8 percent for Claude Opus 4.6 Max, and on CoWorkBench, with 70.7 percent compared with 68.2 percent. The company created both tests itself, described each in a single sentence, and published neither the task list nor the number of samples. As a result, no one can independently reproduce the scores. Alibaba also acknowledges that it fixed problematic tasks in SWE-bench Pro and retested all competing models on the modified dataset, and that it changed several incorrect reference answers in the MathVision and CharXiv tests. This makes sense from an engineering perspective, but the resulting figures cannot be directly compared with scores measured on the original versions of the tests.

For general reasoning, the picture is more mixed than it is for coding. Qwen3.8-27B achieves 89.2 percent on GPQA Diamond, compared with 91.3 percent for Claude Opus 4.6 Max, and only 30.8 percent on Humanity's Last Exam, compared with 40 percent. Conversely, it leads on LiveCodeBench v6 with 90.3 percent versus 88.8 percent, and on IFBench with 79.5 percent versus 62.5 percent. It maintains its largest leads in image and screen interaction: 84.3 percent on OSWorld-Verified versus 72.7 percent, 81.9 percent on AndroidWorld versus 62 percent, and 90 percent on MathVision versus 65.5 percent.

Competing with Meta’s Muse Glimmer

The most direct comparison is with Muse Glimmer, Meta’s 30-billion-parameter open model designed for autonomous operation. It targets the same use case: an always-on assistant running on local hardware. Across the four tests for which Alibaba provides figures for both models, Qwen3.8-27B wins every time: Terminal Bench 2.1 (73% versus 51.7%); SWE-bench Pro (61.7% versus 51.2%); IFBench (79.5% versus 77%); and GPQA Diamond (89.2% versus 83.5%). These are figures that Alibaba itself measured for a competing model, so it is reasonable to treat the margins with caution. On the other hand, the company published both the testing environment and the generation parameters, which is often not the case with similar tables.

Within its own product family, the 27B model is the realistically deployable option. Qwen3.8-Max is a Mixture of Experts model with 2.4 trillion parameters in total and approximately 95 billion active parameters. Alibaba launched it on August 3 and released its weights on August 12, but they are text-only. Qwen3.8-27B is about one hundred times smaller, retains support for images and video, and represents a model that most teams will actually be able to run locally.

Why Alibaba releases open models

There is currently fierce competition for position in the open-weight market. Last week, Meta announced that it would open its most powerful model and add versions tailored for laptops. Alibaba then launched its 27B model and the Max model’s weights into this competitive environment. According to CNBC, this is no coincidence. A model running on a user’s own machine responds faster and keeps data local, which many users consider a crucial privacy advantage. Alongside Alibaba, DeepSeek and Moonshot also remain strong players in the open category. Although Meta started this entire trend with its Llama family, Chinese laboratories are now overtaking it in many respects.

The difference in adoption is clearly visible in the statistics. According to CNBC, the Hugging Face platform reports that developers have already created 151,448 derivative versions based on Qwen models, 2.6 times more than for competing models from Meta. Nick Patience, chief AI analyst at Futurum Group, told CNBC that Alibaba has built Qwen into the most trusted non-American model family. This allows it to establish key partnerships with hardware manufacturers both in China and across the global open-source community. Alibaba itself says that the 27B model matches the performance of a model ten times larger and highlights its benefits for programming, scientific research, office work, and complex autonomous tasks.

Category:AI
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

EU Aims to Catch Up With US and China, Launches Tender for Seven AI GigafactoriesEU Aims to Catch Up With US and China, Launches Tender for Seven AI Gigafactories
Brussels plans to build up to seven massive computing centers for artificial intelligence. Public funding is intended to attract private investment and help Europe catch up with the US and China.
6 min read
24. 8. 2026
Meta Pays Microsoft Hundreds of Millions of Dollars a Year for Rival AI ModelsMeta Pays Microsoft Hundreds of Millions of Dollars a Year for Rival AI Models
Meta is building its own AI infrastructure, but it also uses rival models through Microsoft Azure for development. It pays hundreds of millions of dollars a year to access them.
4 min read
24. 8. 2026
Anthropic Eyes October IPO, Likely to Surpass SpaceX’s Record DebutAnthropic Eyes October IPO, Likely to Surpass SpaceX’s Record Debut
The maker of Claude models is preparing a historic offering at a valuation of up to $2 trillion. But its soaring revenue growth comes with massive losses and AI development costs.
3 min read
24. 8. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok