Microsoft Introduces MAI AI Models: A Step Toward Independence from Third-Party AI Models
On Thursday, Microsoft unveiled two new large language models (LLMs) that it developed over the past year. These models, known as MAI, are intended to reduce the company's reliance on OpenAI's artificial intelligence to power Copilot products. According to a blog post by Mustafa Suleyman, Microsoft's head of consumer AI, the company will begin replacing its existing models with these new versions for Copilot features in the coming weeks. Microsoft also confirmed an earlier report by The Information that it plans to offer these models to developers in a similar way to how OpenAI sells models such as GPT-5.
One of the new models responds to voice commands and replies using speech, similar to models developed by OpenAI and Google. This approach enables more natural interaction, where the model not only processes text but also speaks back to users. Mustafa Suleyman leads the team that has been training the MAI models since May 2024, as first reported by The Information. During development, they encountered technical issues, including attempts to replicate OpenAI's models, while Microsoft has exclusive access to OpenAI's intellectual property.
A Shift Toward Self-Reliance
These models follow Microsoft's investments of more than $10 billion in OpenAI (approximately CZK 225 billion at the current exchange rate), but recent tensions over intellectual property and revenue sharing have prompted the company to seek greater self-reliance. Microsoft established its AI division roughly 18 months ago with the aim of building its own models and reducing its dependence on external providers such as OpenAI. The MAI models are designed to deliver competitive performance, cost efficiency, and greater control over integration into products such as Copilot, which is deeply embedded in Windows and Office.
Microsoft plans to develop a range of specialized models for different user needs, with further advances expected soon. In healthcare, it is developing models such as MAI-DxO for medical diagnostics, with the goal of achieving superintelligence and providing cost-effective support in clinical practice.
Technical Details of the Models
MAI-1-preview is a foundational text model trained on approximately 15,000 Nvidia H-100 graphics processing units, significantly fewer than the 100,000 or more used by some competitors. It is estimated to have around 500 billion parameters, putting it in direct competition with models such as OpenAI's GPT-4. The training focused on data efficiency, using open techniques to maximize output from fewer resources. The model is currently undergoing public testing and community evaluation and is expected to power future versions of Copilot.
The second model, MAI-Voice-1, is a highly efficient speech-generation model capable of producing one minute of audio in less than a second on a single graphics processing unit. It provides high-quality, expressive audio for both single-speaker and multi-speaker scenarios. It is available in preview in applications such as Copilot Daily, Podcasts, and Copilot Labs. Both models are Microsoft's first in-house AI models and are intended to strengthen its position in artificial intelligence.
Future Outlook and Impact
Microsoft confirms that these models are part of a strategy focused on greater independence in AI, which could affect the dynamics of its partnership with OpenAI and the broader market. With public testing underway and plans to expand access for developers, the MAI models are expected to bring innovation to everyday applications such as voice assistance and text processing. Microsoft is responding to competitive pressure from OpenAI and Google while focusing on practical and efficient solutions for users.
Source: www.theinformation.com/



