What Does Mistral 3 Bring?
French company Mistral AI has just introduced its new family of models called Mistral 3. This family includes the large Mistral Large 3 model and nine smaller models under the name Ministral 3. All these models are open-weight, meaning their weights are publicly available under the Apache 2.0 license. This allows anyone to download and use them. The models support multimodal capabilities, such as processing both text and images, and work in many languages, including Czech.
Mistral Large 3 is built on a mixture-of-experts (MoE) architecture, with 41 billion active parameters and 675 billion parameters in total. Training was carried out on 3,000 NVIDIA H200 graphics processors. This model achieves high efficiency by activating only the necessary parts of the network for each token. It has a context window of 256,000 tokens, allowing it to process long documents. According to benchmarks, it ranks second among open-weight models without advanced reasoning on the LMArena leaderboard.
The smaller Ministral 3 models come in sizes of 3 billion, 8 billion, and 14 billion parameters. Each size has three variants: base (basic pretrained), instruct (optimized for conversations and assistants), and reasoning (for complex logical tasks). These models achieve high accuracy, such as 85% on the AIME 2025 test for the 14-billion-parameter variant. They are designed to run on a single GPU, making them suitable for devices without a constant internet connection, such as laptops, robots, or drones.
Benefits for Businesses and Developers
According to Guillaume Lample, co-founder and chief scientist at Mistral AI, companies often start with large closed models but then switch to smaller, customized versions due to lower costs and greater speed. Mistral 3 offers exactly that—the ability to customize models for specific tasks, such as document analysis, code generation, or workflow automation. Ministral 3 generates fewer tokens than comparable models, reducing energy consumption and increasing speed.
Mistral AI is working with companies such as Helsing on drone models that combine vision, language, and actions, and with Stellantis on an in-car assistant. It is also working with the HTX agency in Singapore on models for robots, cybersecurity, and fire protection. These models run offline, which is crucial in places with limited connectivity, such as remote areas or student projects.
Optimization and Availability
Mistral AI is working closely with Nvidia to optimize the models for its hardware. Mistral Large 3 achieves a tenfold speedup on the GB200 NVL72 system compared to the previous-generation NVIDIA H200. This means lower costs per token and greater energy efficiency. The smaller Ministral 3 models are tailored for edge devices such as Nvidia Spark, RTX computers, laptops, or Jetson devices.
Mistral is also working with vLLM and Red Hat on compressed formats such as NVFP4, enabling Large 3 to run on a single node with 8x A100 or H100. The models are available on platforms including Mistral AI Studio, Amazon Bedrock, Azure Foundry, Hugging Face, Modal, IBM WatsonX, OpenRouter, Fireworks, Unsloth AI, and Together AI. Nvidia NIM and AWS SageMaker will be added soon.
Mistral AI also offers services for custom model training on specific data, helping companies create tailored solutions. The models support frameworks such as TensorRT-LLM, SGLang, and vLLM for efficient deployment from the cloud to the edge.
Multimodal Capabilities
All models in the Mistral 3 family process not only text but also images, making them suitable for tasks such as image description or combined analysis. They support more than 40 languages, which is an advantage for global use. For example, Large 3 is the first open-weight model to combine multimodal and multilingual capabilities in a single package, comparable to models such as Meta's Llama 3 or Alibaba's Qwen3-Omni.
The smaller Ministral 3 models have a context window of 128,000 to 256,000 tokens, allowing them to process long conversations or documents. The reasoning variants achieve high accuracy in tests such as GPQA, where they outperform comparable models in their class.
This family of models is designed to be accessible to everyone—from developers to large companies. Mistral AI emphasizes that AI should not be controlled by only a few large laboratories, which is why it is releasing everything openly.
Additional sources: techcrunch.com



