Microsoft Unveils Maia 200: New AI Chip to Rival Nvidia, Google, and Amazon

Microsoft Unveils Maia 200: New AI Chip to Rival Nvidia, Google, and Amazon

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
28. 1. 2026
5 minutes reading
Microsoft Unveils Maia 200: New AI Chip to Rival Nvidia, Google, and Amazon

Microsoft has just unveiled its new artificial intelligence chip called Maia 200, which aims to dramatically improve the economics of AI token generation and reduce the technology giant's dependence on Nvidia chips. It is the second chip developed in-house by Microsoft, with ambitions to compete not only with the dominant Nvidia but also with Google's and Amazon's proprietary solutions.

Impressive technical specifications

Maia 200 is a true technological gem. The chip is manufactured using Taiwanese company TSMC's cutting-edge 3-nanometer process and contains more than 140 billion transistors. Each Maia 200 chip can deliver more than 10 petaflops of performance at 4-bit precision (FP4) and approximately 5 petaflops at 8-bit precision (FP8), all within a power envelope of 750 watts.

For comparison, Microsoft claims that Maia 200 offers three times the FP4 performance of third-generation Amazon Trainium chips and higher FP8 performance than the seventh generation of Google's TPUs. However, it is not just about raw performance. Microsoft emphasizes that Maia 200 is the most efficient inference system the company has ever deployed, offering a 30% better price-performance ratio than the latest generation of hardware in its current portfolio.

Cutting-edge memory and data throughput

One of Maia 200's key advantages is its redesigned memory subsystem. The chip features 216 GB of HBM3e memory with bandwidth of 7 TB per second and 272 MB of on-chip SRAM. This memory system is designed specifically for low-precision data types and includes a specialized DMA engine and NoC fabric for high-speed data transfer, significantly increasing token throughput.

Scott Guthrie, Microsoft's executive vice president responsible for Azure and cloud solutions, explains: "Maia 200 is an accelerator built on TSMC's 3nm process with native FP8/FP4 tensor cores. In practical terms, a single Maia 200 node can easily run today's largest models with sufficient headroom for even larger models in the future."

Maia 200 chip compared with competitors
Maia 200 chip compared with competitors.

Next-generation network architecture

At the system level, Maia 200 introduces an innovative two-tier network design built on standard Ethernet. A custom transport layer and tightly integrated network interface card unlock performance, strong reliability, and significant cost advantages without the need to rely on proprietary network fabrics.

Each accelerator offers 2.8 TB per second of bidirectional dedicated bandwidth and predictable, high-performance collective operations across clusters of up to 6,144 accelerators. In each tray, four Maia accelerators are fully interconnected through direct, non-switched links, keeping high-speed communication local for optimal inference efficiency.

Where and how Maia 200 is used

Maia 200 has already been deployed at Microsoft's data center in the US Central region near Des Moines, Iowa, with another deployment in the US West 3 region near Phoenix, Arizona, on the way and more regions to follow. The chip will serve multiple models, including OpenAI's latest GPT-5.2 models, delivering a price-performance advantage for Microsoft Foundry and Microsoft 365 Copilot.

The Microsoft Superintelligence team will use Maia 200 for synthetic data generation and reinforcement learning to improve next-generation models developed in-house. For synthetic data pipeline use cases, Maia 200's unique design helps accelerate the rate at which high-quality, domain-specific data can be generated and filtered.

Competition in custom AI chips

Microsoft is not the only technology giant trying to reduce its dependence on Nvidia. Google has been using its own TPUs (Tensor Processing Units) for years. They are not sold as standalone chips but as computing power available through the cloud. Amazon has its own AI accelerator chip, Trainium, whose latest version, Trainium3, was launched in December.

According to Daniel Howley, technology editor at Yahoo Finance, while cloud companies such as Google, Amazon, and Microsoft are developing their own AI chips, they are unlikely to pose a serious threat to Nvidia's leadership. Experts say that while cloud companies' AI chips may work well for their own services, this will probably not translate as easily to smaller third-party customers. Nvidia's chips are also highly valued because they are designed for general-purpose use, allowing companies to use them for a wide range of applications and services.

Rapid development and deployment

Microsoft emphasizes that a key principle of its silicon development program is to validate as much of the end-to-end system as possible before the final silicon becomes available. A sophisticated pre-silicon environment guided the Maia 200 architecture from its earliest stages, modeling the computational and communication patterns of large language models with high fidelity.

Thanks to these investments, AI models were running on Maia 200 silicon within days of the first packaged components arriving. The time from first silicon to the first rack deployment in a data center was reduced by more than half compared with similar AI infrastructure programs.

Microsoft clearly states that the era of large-scale artificial intelligence is only just beginning and that infrastructure will define what is possible. The Maia AI accelerator program is designed as a multigenerational effort. While the company deploys Maia 200 across its global infrastructure, it is already designing future generations and expects each generation to continually set new standards for what is possible while delivering ever-better performance and efficiency for the most important AI workloads.

SDK for developers

Microsoft is inviting developers, AI startups, and academics to begin exploring early model and workload optimization with the new Maia 200 software development kit (SDK). The SDK includes the Triton Compiler, support for PyTorch, low-level programming in NPL, and a Maia simulator and cost calculator for optimizing efficiency earlier in the code lifecycle.

According to Lucas Ropek of TechCrunch, inference refers to the computational process of running a model, as opposed to the computing power required to train it. As AI companies mature, inference costs have become an increasingly important part of their overall operating expenses, leading to renewed interest in ways to optimize the process.

Maia 200 integrates seamlessly with Azure, and Microsoft is now offering a preview of the Maia SDK with a complete suite of tools for building and optimizing models for Maia 200. This gives developers fine-grained control when needed while enabling models to be easily ported across heterogeneous hardware accelerators.

Category:AI
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

Altman Announced the Singularity Days After His Models Escaped the Lab on Their OwnAltman Announced the Singularity Days After His Models Escaped the Lab on Their Own
OpenAI chief Sam Altman declared on the Relentless podcast that humanity has already entered the singularity. “We’re like, in the singularity now,” he said verbatim. For decades, the term belonged more to science-fiction literature
6 min read
28. 7. 2026
AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.
Since April, Australian radio has been playing a dance remake of Madonna’s hit Like a Prayer on repeat. Released by Queensland DJ Josh Fawaz, it tops the radio airplay chart and has 35 million Spotify streams.
6 min read
28. 7. 2026
Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?
A first-person shooter that runs directly in the browser, with its own physics and eleven separate code modules. Around 55,000 lines in total, split across eleven subsystems and built on Thr
4 min read
28. 7. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok