For most of its history, Microsoft was a software company. Windows, Office, and Azure ran on silicon designed by Intel, AMD, and later Nvidia. In the AI era, that arrangement has become one of the most expensive dependencies in tech, and Microsoft is moving to change it.
The cost problem
Running large language models at scale is brutally expensive. Every Copilot query and every Azure OpenAI call consumes accelerator time. When Microsoft first introduced Maia at Ignite in 2023, the goal was to cut those costs and power its Copilot services, with the chip optimized for the shared foundation models behind nearly all of Microsoft’s AI products.
At Microsoft’s scale, even a small efficiency gain compounds into huge savings. Buying general-purpose GPUs at market prices, with healthy vendor margins, makes that harder.

Breaking the Nvidia dependency
Nvidia still dominates AI hardware, with roughly 70 percent of the AI chip market according to one recent analysis. That concentration creates three risks for any hyperscaler:
- Supply risk: If Nvidia can’t deliver, Azure’s growth stalls.
- Pricing risk: A near-monopoly supplier sets the terms.
- Roadmap risk: Microsoft’s product plans depend on someone else’s release schedule.
Custom silicon gives Microsoft more control over all three. It’s also not alone: Google has its TPUs and Amazon has Trainium, and all three are designing their own hardware to reduce reliance on Nvidia and lower operating costs.
Maia 200: where things stand now
The second-generation Maia 200 shows how serious the effort has become. Bloomberg reported that it is built by TSMC and was rolling out first to Microsoft’s Iowa data centers, with the Phoenix area next. The specs are strong: a 3nm process, over 10 petaFLOPS at 4-bit precision, about 5 petaFLOPS at 8-bit, 216GB of HBM3e memory, and 272MB of on-die SRAM.
Microsoft claims three times the FP4 performance of Amazon’s third-generation Trainium and better FP8 performance than Google’s seventh-generation TPU, with 30% better performance per dollar than its previous hardware. It supports OpenAI’s GPT-5.2 models and Microsoft 365 Copilot.
It’s also moving into real production. At Build 2026, Microsoft said Maia 200 is live in Iowa and Arizona, with Italy, Australia, and South Korea planned next.

Cobalt: the quiet half of the plan
Maia gets the headlines, but the Arm-based Cobalt CPU matters just as much. Maia handles AI acceleration, while Cobalt takes on general-purpose compute in Azure. Cobalt 200 delivers a 50 percent performance jump over Cobalt 100, and Azure VMs based on it are in early access preview. Owning both the accelerator and the CPU lets Microsoft optimize the full stack, from chip to rack to cooling to software.
The software moat
Hardware alone doesn’t win. Nvidia’s real advantage is CUDA, the ecosystem developers already know. Microsoft’s answer is a Maia SDK with PyTorch integration, a Triton compiler, and optimized kernel libraries, so developers can port models across accelerators. Lowering the switching cost matters as much as raw performance.
It hasn’t been smooth
The project has had setbacks. Maia 200 reportedly slipped about six months in mid-2025 because of design changes and staff departures. Competing with established chip makers’ development cycles is hard even for a company of Microsoft’s size, and Nvidia isn’t standing still.
The bigger picture
Microsoft isn’t trying to replace Nvidia overnight. The aim is leverage and efficiency: run its own workloads, like Copilot and OpenAI models, on cheaper in-house silicon, and use Nvidia GPUs where they make the most sense. There’s also a possible new revenue angle, since reports say Anthropic may rent Maia 200 servers to run Claude models. That would be the chip’s first major external use.
In short, Microsoft is building its own AI chips for the same reasons it once built its own software platforms: control, cost, and independence. In a market where compute is the scarcest resource, owning more of the stack is a strategic necessity.
SiliconeUpdate.com is a technology news platform that publishes updates and informational content related to silicon technology, software, artificial intelligence, and emerging technologies.
All articles published on this platform are attributed to SiliconeUpdate.com instead of individual authors. Content is presented in a neutral, informational format without personal opinions.
—
Content Publishing
SiliconeUpdate.com publishes news and updates based on publicly available information, official announcements, and industry developments. The focus is on clarity, relevance, and timely reporting.
—
Editorial Control
All editorial decisions, updates, and content management are handled at the platform level. No individual human or AI identity is presented as the author of articles.
—
Contact
For editorial communication or general queries, contact:
Email: neemasharma@gmail.com