Microsoft plans to publicly show its next custom AI chip, the Maia 300, as early as next month, The Information reported August 10, attributing the timeline to unnamed people close to the effort. It follows the original Maia chip that Microsoft first introduced in November 2023. A public reveal would be Microsoft’s clearest move yet toward running its own cloud AI infrastructure on less Nvidia silicon, though the reveal itself does not tell us how much of that infrastructure will actually change hands.

Microsoft has already secured production capacity at Taiwan Semiconductor Manufacturing Co. (TSMC) covering over 300,000 Maia 300 units, with deliveries slated for 2027, according to the report. Locking in that volume two years ahead of shipment suggests Microsoft is treating this generation as a production part rather than a one-off demo.

The original Maia chip packed 105 billion transistors on a 5-nanometer node and was built to handle both training and serving of large models, but it stayed mostly an internal tool rather than a product customers could buy. Its rollout lagged Microsoft’s own internal targets and moved slower than competing programs: Google’s TPU v5p already ships broadly through the cloud, and Amazon’s Trainium2 is scaling for large training runs.

Patrick Moorhead, an analyst at Moor Insights & Strategy, said Microsoft’s approach has differed from its rivals. “Microsoft has been more cautious, integrating Maia deeply into Azure’s fabric before exposing it broadly,” Moorhead said in a recent interview. “The Maia 300 could be the moment that changes.”

Microsoft is also trying to win over big cloud customers to run production workloads on the Maia 300, Anthropic among them, according to the report. Anthropic buys substantial Nvidia capacity today and trains and serves its Claude models on it, so any meaningful shift would be among the first visible defections by a frontier lab away from Nvidia’s stack. Nvidia’s CUDA software layer remains the default for a reason: porting production models to new silicon is slow, expensive engineering work, and Microsoft has not yet shown it can offer a tooling stack mature enough to make that trade worthwhile for an outside customer.

Every major hyperscaler now runs a custom silicon program pointed at the same target. Google has TPUs, Amazon has Trainium, and Microsoft has Maia. None of the three has meaningfully displaced Nvidia inside its own production stack, let alone anyone else’s. That is the context a September unveiling sits inside: the existence of a new chip generation is not evidence of a shift in where the compute actually runs.

The more useful number is one Microsoft has never published for any Maia generation: what share of its own internal training and inference Maia is actually allowed to carry, as opposed to what share stays on Nvidia GPUs because CUDA compatibility, latency guarantees, or customer contracts require it. A reveal event does not answer that question, and Microsoft has given no indication it intends to.

Operators watching this story for signal should track whether Microsoft or Azure ever discloses a workload-share figure once the Maia 300 ships in 2027, not whether the September event lands well. Until that figure exists, the safer assumption is that Nvidia keeps the large majority of Microsoft’s production AI compute.

Inside AI News reported this story on August 10, 2026, drawing on The Information’s account of Microsoft’s plans.