DeepSeek V3.1: The Most Powerful Open AI Model of 2025?

DeepSeek V3.1: The Most Powerful Open AI Model of 2025?

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
25. 8. 2025
4 minutes reading
DeepSeek V3.1: The Most Powerful Open AI Model of 2025?

DeepSeek V3.1: The Most Powerful Open AI Model of 2025?

Whether you are a developer looking for an affordable tool for complex tasks or simply a curious artificial intelligence enthusiast, the new DeepSeek V3.1 model will make you stop and think. This Chinese colossus with 685 billion parameters delivers performance comparable to the most expensive proprietary systems, but at a fraction of the price. Let’s take a look at what makes it so exceptional—and why speculation is swirling around the future of the entire model family.

Technical Specifications

DeepSeek V3.1 is a massive large language model (LLM) with a total of 685 billion parameters, activating approximately 37 billion of them during inference thanks to its Mixture-of-Experts (MoE) architecture. This means the model is efficient and does not consume excessive computing power—only a fraction of its total capacity is used per token. A major improvement over the previous V3 version is the expanded context window of 128,000 tokens, equivalent to processing information from roughly a 300-page book in a single session. This model underwent two-stage training on long contexts: first on 630 billion tokens to expand the context to 32k, followed by another 209 billion for the full 128k.

DeepSeek, founded by entrepreneur Liang Wenfeng as a side project of his quantitative trading firm, announced this update with only a brief message in one of its WeChat groups. No major press conference, no fanfare on social networks such as X—just a quiet release that immediately attracted attention. The model supports various precision formats, including BF16 (Brain Float 16, a 16-bit format from Google), F8_E4M3 (an 8-bit floating-point format), and F32 (standard 32-bit precision), enabling optimization for different hardware. It also uses FP8 microscaling for efficient large-scale inference, accelerating responses and reducing costs.

Reasoning and Fast Answers in One Package

What makes DeepSeek V3.1 truly interesting is its hybrid architecture. The model can seamlessly switch between a "thinking" mode (chain-of-thought reasoning, similar to the previous R1) and a "non-thinking" mode for direct, fast answers. All you need to do is adjust the prompt (input instruction) or use the "reasoning enabled" boolean switch in the API. This approach covers a wide range of tasks—from ordinary chat and complex logical problems to the use of tools such as web search or coding.

Unlike previous versions, which required separate models for conversations (such as DeepSeek-V2) and reasoning (DeepSeek-R1), everything now runs within a single system. The model is trained to natively support programming and agents (autonomous systems), making it ideal for developers. According to tests on the Aider Polyglot benchmark, it even outperforms Claude 4 Opus on complex multilingual programming tasks, achieving a success rate of 71.6% in Aider tests. And it does all this with greater token efficiency and faster responses than R1.

Cheaper Than the Competition, but Just as Powerful

DeepSeek V3.1 achieves performance comparable to proprietary models such as Claude Opus 4, scoring 71.6% on the SWE-bench benchmark. In mathematical tests such as AIME 2024 or MATH 500, it surpasses many rivals thanks to its strong logical reasoning abilities—for example, it solves problems involving a bouncing ball inside a rotating hexagon. It is the best non-TTC (non-tool-tuned coding) model in its class for programming.

And now for the most compelling part: the price. Through the API, it costs USD 0.56 (approximately CZK 12.32) per million input tokens and USD 2.19 (about CZK 48.18) per million output tokens. That is around 68 times cheaper than Claude Opus, where an equivalent task costs roughly USD 70 (around CZK 1,540). Training the previous V3 version cost only USD 5.6 million (approximately CZK 123.2 million), a fraction of the cost incurred by American laboratories. The model is available on Hugging Face under the MIT license for commercial use, through the web interface at chat.deepseek.com with the “DeepThink” feature, or via the OpenRouter API.

V3.1 benchmark

R2 Speculation: Where Did R1 Go, and What Comes Next?

Interestingly, DeepSeek removed all references to the R1 model from the “deep think” feature in its chatbot. This sparked speculation about the fate of the anticipated R2 successor. The company, which released V3 in December and R1 in January, is now delivering only incremental updates, while competitors such as OpenAI and Anthropic continue to churn out new models. According to an article by Ben Jiang in the South China Morning Post, this suggests a possible shift in the company’s research focus. Patrick Zandl notes on marigold.cz that the quiet launch of V3.1 could be a strategy focused on community validation rather than a major announcement.

The geopolitical aspects cannot be overlooked—American companies may hesitate because of tensions between the United States and China, even though the model is open. In addition, its 700 GB size makes local hosting challenging, so most users will rely on cloud APIs. Nevertheless, the Hugging Face community quickly embraced the model, with more than 80,000 followers shortly after its release.

Category:AI
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

Altman Announced the Singularity Days After His Models Escaped the Lab on Their OwnAltman Announced the Singularity Days After His Models Escaped the Lab on Their Own
OpenAI chief Sam Altman declared on the Relentless podcast that humanity has already entered the singularity. “We’re like, in the singularity now,” he said verbatim. For decades, the term belonged more to science-fiction literature
6 min read
28. 7. 2026
AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.
Since April, Australian radio has been playing a dance remake of Madonna’s hit Like a Prayer on repeat. Released by Queensland DJ Josh Fawaz, it tops the radio airplay chart and has 35 million Spotify streams.
6 min read
28. 7. 2026
Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?
A first-person shooter that runs directly in the browser, with its own physics and eleven separate code modules. Around 55,000 lines in total, split across eleven subsystems and built on Thr
4 min read
28. 7. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok