Microsoft has just unveiled its new artificial intelligence chip called Maia 200, which aims to dramatically improve the economics of AI token generation and reduce the technology giant's dependence on Nvidia chips. It is the second chip developed in-house by Microsoft, with ambitions to compete not only with the dominant Nvidia but also with Google's and Amazon's proprietary solutions.
Impressive technical specifications
Maia 200 is a true technological gem. The chip is manufactured using Taiwanese company TSMC's cutting-edge 3-nanometer process and contains more than 140 billion transistors. Each Maia 200 chip can deliver more than 10 petaflops of performance at 4-bit precision (FP4) and approximately 5 petaflops at 8-bit precision (FP8), all within a power envelope of 750 watts.
For comparison, Microsoft claims that Maia 200 offers three times the FP4 performance of third-generation Amazon Trainium chips and higher FP8 performance than the seventh generation of Google's TPUs. However, it is not just about raw performance. Microsoft emphasizes that Maia 200 is the most efficient inference system the company has ever deployed, offering a 30% better price-performance ratio than the latest generation of hardware in its current portfolio.
Cutting-edge memory and data throughput
One of Maia 200's key advantages is its redesigned memory subsystem. The chip features 216 GB of HBM3e memory with bandwidth of 7 TB per second and 272 MB of on-chip SRAM. This memory system is designed specifically for low-precision data types and includes a specialized DMA engine and NoC fabric for high-speed data transfer, significantly increasing token throughput.
Scott Guthrie, Microsoft's executive vice president responsible for Azure and cloud solutions, explains: "Maia 200 is an accelerator built on TSMC's 3nm process with native FP8/FP4 tensor cores. In practical terms, a single Maia 200 node can easily run today's largest models with sufficient headroom for even larger models in the future."
Next-generation network architecture
At the system level, Maia 200 introduces an innovative two-tier network design built on standard Ethernet. A custom transport layer and tightly integrated network interface card unlock performance, strong reliability, and significant cost advantages without the need to rely on proprietary network fabrics.
Each accelerator offers 2.8 TB per second of bidirectional dedicated bandwidth and predictable, high-performance collective operations across clusters of up to 6,144 accelerators. In each tray, four Maia accelerators are fully interconnected through direct, non-switched links, keeping high-speed communication local for optimal inference efficiency.
Where and how Maia 200 is used
Maia 200 has already been deployed at Microsoft's data center in the US Central region near Des Moines, Iowa, with another deployment in the US West 3 region near Phoenix, Arizona, on the way and more regions to follow. The chip will serve multiple models, including OpenAI's latest GPT-5.2 models, delivering a price-performance advantage for Microsoft Foundry and Microsoft 365 Copilot.
The Microsoft Superintelligence team will use Maia 200 for synthetic data generation and reinforcement learning to improve next-generation models developed in-house. For synthetic data pipeline use cases, Maia 200's unique design helps accelerate the rate at which high-quality, domain-specific data can be generated and filtered.
Competition in custom AI chips
Microsoft is not the only technology giant trying to reduce its dependence on Nvidia. Google has been using its own TPUs (Tensor Processing Units) for years. They are not sold as standalone chips but as computing power available through the cloud. Amazon has its own AI accelerator chip, Trainium, whose latest version, Trainium3, was launched in December.
According to Daniel Howley, technology editor at Yahoo Finance, while cloud companies such as Google, Amazon, and Microsoft are developing their own AI chips, they are unlikely to pose a serious threat to Nvidia's leadership. Experts say that while cloud companies' AI chips may work well for their own services, this will probably not translate as easily to smaller third-party customers. Nvidia's chips are also highly valued because they are designed for general-purpose use, allowing companies to use them for a wide range of applications and services.
Rapid development and deployment
Microsoft emphasizes that a key principle of its silicon development program is to validate as much of the end-to-end system as possible before the final silicon becomes available. A sophisticated pre-silicon environment guided the Maia 200 architecture from its earliest stages, modeling the computational and communication patterns of large language models with high fidelity.
Thanks to these investments, AI models were running on Maia 200 silicon within days of the first packaged components arriving. The time from first silicon to the first rack deployment in a data center was reduced by more than half compared with similar AI infrastructure programs.
Microsoft clearly states that the era of large-scale artificial intelligence is only just beginning and that infrastructure will define what is possible. The Maia AI accelerator program is designed as a multigenerational effort. While the company deploys Maia 200 across its global infrastructure, it is already designing future generations and expects each generation to continually set new standards for what is possible while delivering ever-better performance and efficiency for the most important AI workloads.
SDK for developers
Microsoft is inviting developers, AI startups, and academics to begin exploring early model and workload optimization with the new Maia 200 software development kit (SDK). The SDK includes the Triton Compiler, support for PyTorch, low-level programming in NPL, and a Maia simulator and cost calculator for optimizing efficiency earlier in the code lifecycle.
According to Lucas Ropek of TechCrunch, inference refers to the computational process of running a model, as opposed to the computing power required to train it. As AI companies mature, inference costs have become an increasingly important part of their overall operating expenses, leading to renewed interest in ways to optimize the process.
Maia 200 integrates seamlessly with Azure, and Microsoft is now offering a preview of the Maia SDK with a complete suite of tools for building and optimizing models for Maia 200. This gives developers fine-grained control when needed while enabling models to be easily ported across heterogeneous hardware accelerators.



