In March, Microsoft released MAI-Image-2, its second proprietary text-to-image generation model. The model almost immediately ranked third on the global Arena.ai leaderboard, just behind Google and OpenAI.
Microsoft was long perceived primarily as OpenAI’s largest investor. Copilot, Bing Image Creator, and the entire visual side of its products were built on external foundations. MAI-Image-2 changes all of that. The Microsoft AI Superintelligence team, led by Mustafa Suleyman, decided to forge its own path, and the result surprised even skeptics. The model was developed in close collaboration with photographers, designers, and visual creators. It was designed with real-world use in mind. And it shows.
What can the new MAI-Image-2 do?
Three areas in which the model excels: photorealism, text generation within images, and complex scenes.
Photorealism is not just a marketing buzzword. Microsoft has solved a problem that has plagued generative AI from the beginning: so-called “plastic” skin. Collaboration with professional photographers helped fine-tune the algorithms for realistic rendering of textures and lighting physics. Portraits from MAI-Image-2 look like they came from a studio, not a computer.
Then there is typography. A 115-point increase in the text accuracy benchmark is a figure that graphic designers will appreciate immediately. Infographics, banners, and posters with legible text directly in the image. And crucially for Czech users: the model natively supports Czech diacritics. Carons, accents, no garbling. Until now, this was the number one pain point when deploying AI in Czech marketing.

The third major strength is complex scenes. Surrealist concepts, cinematic compositions, detailed worlds. MAI-Image-2 can handle them all.
Technically speaking: Flow matching and up to 50 billion parameters
Under the hood is an architecture based on the flow-matching diffusion method, which enables a smoother transition from digital noise to a clean image than conventional models. The estimated parameter count is 10 to 50 billion, the output resolution is currently fixed at 1024x1024 pixels, and the context window reaches 32,000 tokens, making it possible to process even highly detailed and complex prompts without losing coherence.
The free version within Copilot has its limitations: a 30-second delay between generations, a limit of 15 images per day, and only a square 1:1 format. Its full potential is unlocked in Microsoft Foundry, where companies can fine-tune the model using their own brand data.

Where can you find MAI-Image-2 and how much does it cost?
The model is available through MAI Playground on microsoft.ai and is gradually being rolled out to Copilot and Bing Image Creator, though it is not yet available in the Czech Republic. Enterprise API access is currently available to selected customers. Broader access through Microsoft Foundry is coming soon. In terms of pricing, individual creators will pay approximately CZK 3,499 per year for Copilot Pro, with 100 priority generations per day. Small and medium-sized businesses can obtain Microsoft 365 Copilot for EUR 15.60 per month with annual billing until June 2026.
MAI-Image-2 vs. Midjourney: Who wins?
It depends on what you need.
Midjourney v7 still leads in aesthetic coherence and artistic expression. If you are looking for dreamlike visuals or abstract concepts, it remains the tool for you. But as soon as you need a marketing banner with precise text, an infographic with Czech diacritics, or product photos with natural lighting, MAI-Image-2 currently has no comparable competition.
Generation speed is comparable to Midjourney’s Draft mode, but at full resolution. According to Microsoft, the model is also more energy-efficient than previous generations, aligning with the company’s sustainability goals.
Data security for businesses
Microsoft built MAI-Image-2 on the principle of Tenant Isolation: your company data remains in an isolated environment and is never used to train public models. All inputs and outputs are covered by Enterprise Data Protection.
In addition, every generated image contains C2PA standard metadata and invisible watermarks resistant to cropping and compression. This is a direct response to the upcoming EU AI Act, whose key Article 50 takes effect in August 2026.
Microsoft has also expanded its Customer Copyright Commitment: if a user complies with the built-in safety filters and works under a paid license, Microsoft assumes legal responsibility for potential intellectual property disputes. For Czech companies, this is a safeguard that significantly lowers the barrier to widespread deployment.
Graphic designers are not becoming obsolete, but their role is changing
Deploying MAI-Image-2 in PowerPoint reduces the time needed to create professional presentations from an average of four hours to 45 minutes. In Word, the model acts as a visual assistant that reads the document’s context and automatically suggests relevant illustrations. In Teams, the Visual Canvas feature enables collaborative visual creation directly during a call.
Graphic designers are not becoming obsolete. They are moving away from redrawing slides and starting to define visual strategy. Their role is shifting from manual execution to management.
Mustafa Suleyman predicts that most administrative and creative tasks will be fully automated within 12 to 18 months. Time will tell whether he is right. But MAI-Image-2 is a compelling first step in that direction.



