Why Microsoft Is Building Its Own AI Chips

Why Microsoft Is Building Its Own AI Chips

Share

For most of its history, Microsoft was a software company. Windows, Office, and Azure ran on silicon designed by Intel, AMD, and later Nvidia. In the AI era, that arrangement has become one of the most expensive dependencies in tech, and Microsoft is moving to change it.

The cost problem

Running large language models at scale is brutally expensive. Every Copilot query and every Azure OpenAI call consumes accelerator time. When Microsoft first introduced Maia at Ignite in 2023, the goal was to cut those costs and power its Copilot services, with the chip optimized for the shared foundation models behind nearly all of Microsoft’s AI products.

At Microsoft’s scale, even a small efficiency gain compounds into huge savings. Buying general-purpose GPUs at market prices, with healthy vendor margins, makes that harder.

at Microsoft's scale, even a small efficiency gain compounds into huge savings. Buying general-purpose GPUs at market prices, with healthy vendor margins, makes that harder.

Breaking the Nvidia dependency

Nvidia still dominates AI hardware, with roughly 70 percent of the AI chip market according to one recent analysis. That concentration creates three risks for any hyperscaler:

  • Supply risk: If Nvidia can’t deliver, Azure’s growth stalls.
  • Pricing risk: A near-monopoly supplier sets the terms.
  • Roadmap risk: Microsoft’s product plans depend on someone else’s release schedule.

Custom silicon gives Microsoft more control over all three. It’s also not alone: Google has its TPUs and Amazon has Trainium, and all three are designing their own hardware to reduce reliance on Nvidia and lower operating costs.

Maia 200: where things stand now

The second-generation Maia 200 shows how serious the effort has become. Bloomberg reported that it is built by TSMC and was rolling out first to Microsoft’s Iowa data centers, with the Phoenix area next. The specs are strong: a 3nm process, over 10 petaFLOPS at 4-bit precision, about 5 petaFLOPS at 8-bit, 216GB of HBM3e memory, and 272MB of on-die SRAM.

Microsoft claims three times the FP4 performance of Amazon’s third-generation Trainium and better FP8 performance than Google’s seventh-generation TPU, with 30% better performance per dollar than its previous hardware. It supports OpenAI’s GPT-5.2 models and Microsoft 365 Copilot.

It’s also moving into real production. At Build 2026, Microsoft said Maia 200 is live in Iowa and Arizona, with Italy, Australia, and South Korea planned next.

It's also moving into real production. At Build 2026, Microsoft said Maia 200 is live in Iowa and Arizona, with Italy, Australia, and South Korea planned next.

Cobalt: the quiet half of the plan

Maia gets the headlines, but the Arm-based Cobalt CPU matters just as much. Maia handles AI acceleration, while Cobalt takes on general-purpose compute in Azure. Cobalt 200 delivers a 50 percent performance jump over Cobalt 100, and Azure VMs based on it are in early access preview. Owning both the accelerator and the CPU lets Microsoft optimize the full stack, from chip to rack to cooling to software.

The software moat

Hardware alone doesn’t win. Nvidia’s real advantage is CUDA, the ecosystem developers already know. Microsoft’s answer is a Maia SDK with PyTorch integration, a Triton compiler, and optimized kernel libraries, so developers can port models across accelerators. Lowering the switching cost matters as much as raw performance.

It hasn’t been smooth

The project has had setbacks. Maia 200 reportedly slipped about six months in mid-2025 because of design changes and staff departures. Competing with established chip makers’ development cycles is hard even for a company of Microsoft’s size, and Nvidia isn’t standing still.

The bigger picture

Microsoft isn’t trying to replace Nvidia overnight. The aim is leverage and efficiency: run its own workloads, like Copilot and OpenAI models, on cheaper in-house silicon, and use Nvidia GPUs where they make the most sense. There’s also a possible new revenue angle, since reports say Anthropic may rent Maia 200 servers to run Claude models. That would be the chip’s first major external use.

In short, Microsoft is building its own AI chips for the same reasons it once built its own software platforms: control, cost, and independence. In a market where compute is the scarcest resource, owning more of the stack is a strategic necessity.


Scroll to Top