Microsoft came to custom silicon later than Google or Amazon and for a sharper reason: it had committed to serving OpenAI models at a scale that made its NVIDIA bill a strategic problem rather than a line item.
Maia is explicitly an inference chip. That is the honest place for a first-generation accelerator — training a frontier model demands a mature software stack, while serving one is a more tractable engineering problem with far larger aggregate volume.
Like every hyperscaler part, it is captive. Microsoft designs chips to reduce what it pays for chips, not to enter the market for them.
