Home Companies Microsoft Maia
Cloud-captive silicon Part of Microsoft — NASDAQ: MSFT In production

Microsoft Maia

Maia 200 is the first Microsoft accelerator with enough disclosed detail to assess properly — and it is explicitly an inference chip, not a general-purpose one. It already serves OpenAI's production models inside Azure, which is a stronger validation than most captive silicon can claim.

Key facts

Programme
Maia, Microsoft's custom AI accelerator line
First generation
Maia 100, 2023
Current generation
Maia 200, announced 26 January 2026
Process
TSMC 3nm, over 140 billion transistors
Compute
Over 10 PFLOPS FP4; over 5 PFLOPS FP8
Memory
216GB HBM3e at 7 TB/s, plus 272MB on-chip SRAM
Power
750W TDP
Scale-up
2.8 TB/s per accelerator; collectives across up to 6,144 accelerators
Deployed
US Central (Iowa), with US West 3 (Arizona) next
Workloads
OpenAI GPT-5.2, Microsoft Foundry, Microsoft 365 Copilot
Sold externally
No
Last reviewed
21 August 2026

Overview

Maia is Microsoft's custom AI accelerator line, built for Azure. It is the youngest of the three major hyperscaler programmes — Maia 100 arrived in 2023, where Google had been shipping TPUs since 2015 and AWS had Annapurna Labs from 2015 — and until recently it was also the least substantiated, with far less public technical detail than its peers.

Maia 200 changed that. Announced on 26 January 2026, it is a genuinely detailed disclosure: built on TSMC's 3nm process with more than 140 billion transistors, delivering over 10 petaFLOPS of FP4 and more than 5 petaFLOPS of FP8 compute within a 750W envelope, paired with 216GB of HBM3e at 7 TB/s and 272MB of on-chip SRAM. Each accelerator exposes 2.8 TB/s of dedicated bidirectional scale-up bandwidth, and the design supports collective operations across clusters of up to 6,144 accelerators.

The name matters: Microsoft describes Maia 200 as an inference accelerator, not a general AI chip. Like Google, Microsoft has concluded that inference deserves purpose-built silicon rather than a converged design.

Maia 200 specifications

AttributeFigureWhy it matters
Process and scaleTSMC 3nm, over 140 billion transistorsLeading-edge node, comparable in transistor budget to merchant flagship parts.
ComputeOver 10 PFLOPS FP4; over 5 PFLOPS FP8The emphasis on very low precision is an inference signal — FP4 is useful for serving, not for training stability.
Memory216GB HBM3e at 7 TB/s, plus 272MB on-chip SRAMA large memory capacity per accelerator, which matters for holding big models and long-context KV cache without splitting across devices.
Power750W TDPDisclosed, which is more than most vendors publish, and allows a real performance-per-watt comparison.
Scale-up2.8 TB/s bidirectional per accelerator; collectives across up to 6,144 acceleratorsDesigned for large coherent domains rather than single-node serving.
EconomicsMicrosoft states 30% better performance per dollar than its current hardwareThe stated justification for the programme, though the comparison baseline is not specified in detail.

All figures are Microsoft's own disclosures. Because Maia is not sold, no independent party can benchmark it, and the 30% performance-per-dollar claim cannot be externally verified.

Deployment and workloads

Maia 200 is in production rather than announced-only. Microsoft has deployed it in its US Central region near Des Moines, Iowa, with US West 3 near Phoenix, Arizona next and further regions to follow.

The workloads are the most interesting disclosure. Maia 200 serves OpenAI's GPT-5.2 models, Microsoft Foundry and Microsoft 365 Copilot. That first item is significant: Microsoft is running a frontier partner's production model on its own silicon, which is a stronger validation than internal workloads alone would be.

It also sits in a complicated relationship with OpenAI's own Jalapeño, announced with Broadcom in June 2026 and due to deploy at gigawatt scale from late 2026 through partners including Microsoft. Microsoft is therefore serving OpenAI models on Microsoft-designed silicon while also hosting OpenAI-designed silicon in its data centres. Both companies are attacking the same cost problem from different directions, and the arrangement is cooperative and competitive at once.

Software

Microsoft ships a Maia SDK, currently in preview, with PyTorch integration and a Triton compiler path. The Triton choice is pragmatic: it lets developers write kernels at a higher level than CUDA and target multiple backends, which lowers the porting burden that has stalled every non-Nvidia accelerator.

Because Maia is captive, the software problem is also narrower than it would be for a merchant vendor. Microsoft does not need arbitrary customer code to run well; it needs its own services and a small number of partner models to run well. That is a far more tractable target, and it is the same structural advantage OpenAI has with Jalapeño.

Strengths and key questions

Strengths

What is working

  • Guaranteed demand. Azure AI, Copilot and OpenAI inference provide enormous captive workload with no sales cycle.
  • Real deployment. Maia 200 is serving production traffic in named regions, not sampling to customers.
  • Purpose-built for inference. The FP4 emphasis and memory configuration are coherent choices for serving rather than compromises.
  • Unusual transparency. Publishing transistor count, TDP and memory bandwidth is more than most hyperscalers disclose.
  • Narrow software target. Only Microsoft's own workloads need to run well, which makes the compiler problem tractable.
Key questions

What is unresolved

  • No external verification. The chip is not sold, so every performance figure rests on Microsoft's word.
  • Late start. Maia 200 is Microsoft's second generation against Google's seventh. The compiler and systems experience gap is real.
  • Displacement share. Microsoft has not said what fraction of Azure inference runs on Maia rather than purchased GPUs, which is the number that matters.
  • Overlap with Jalapeño. Microsoft and OpenAI are building separate inference silicon for substantially overlapping workloads.
  • Customer access. Whether Azure customers can select Maia instances directly, or whether it stays behind Microsoft's own services.

What to watch

  • Region rollout. How quickly Maia 200 moves beyond Iowa and Arizona, which indicates real production confidence.
  • Customer-selectable instances. If Azure exposes Maia as a choosable instance type, it becomes a competitive product rather than internal infrastructure.
  • Maia SDK general availability. Movement from preview, and the breadth of model support through the Triton path.
  • Maia 300. Whether Microsoft sustains an annual-ish cadence, which is what separates a programme from an experiment.
  • The OpenAI silicon question. How Maia and Jalapeño divide OpenAI inference inside Microsoft data centres.

Sources

  1. Microsoft — Maia 200: the AI accelerator built for inference (26 January 2026). Primary announcement.
  2. Microsoft Source — Maia 200 introduction. Deployment regions and workload detail.
  3. NAND Research — Research note on the Azure Maia 200. Independent analysis of specifications and positioning.
  4. IN Electronics & Design — Microsoft details the Maia 200. Transistor count, power envelope and scale-up bandwidth.

Maia is captive silicon and is not sold, so all technical and economic claims originate with Microsoft and cannot be independently verified. See the editorial methodology.