Microsoft AI released MAI-Transcribe-2 on Thursday, a speech recognition model priced at $0.10 per hour of audio, according to the company’s own announcement. Microsoft says the model outperforms Gemini 3.5 Transcribe, GPT-Transcribe, Whisper V3-Large and ElevenLabs’ Scribe v2 on both accuracy and speed. It covers 60 languages, separates speakers, stamps every individual word with a time, and lets the caller choose the style of the output. The price is the part worth sitting with: it is roughly a third of what Microsoft charged for the previous version five months ago.
MAI-Transcribe-2 takes the top position on FLEURS over that whole language set, averaging a 5.2% word error rate, Microsoft says, and the company places it at what it calls the Pareto frontier for accuracy against latency on Artificial Analysis, an independent benchmarking firm. Every one of those figures comes from Microsoft’s own release. The company has not published independent third-party audit results alongside the launch, so the accuracy and speed claims should be read as Microsoft’s self-reported numbers until an outside benchmarker replicates them.
VentureBeat, covering the launch separately, reports that the price undercuts OpenAI, Google and ElevenLabs on hourly cost, and reads the release as part of a wider Microsoft habit: standing up its own frontier-grade system in one modality after another, then quietly substituting it into products that had been running on OpenAI’s stack. VentureBeat notes this is the third speech model Microsoft has shipped in five months, following versions in April and June, each adding language coverage and features that speech vendors have historically sold as premium add-ons.
That cadence matters more than any single benchmark line. A lab that ships three iterations of the same product category in five months has settled on an architecture and is scaling data, not still hunting for a working design. Diarization, keyword biasing, and dual verbatim and clean output modes are the kind of features that used to justify a markup from specialty transcription vendors. Bundling them into a base tier priced at a dime signals Microsoft wants transcription treated as infrastructure, not as a discrete purchasing decision.
The bigger shift is what a 10 cent price does to how products get built, not just what it does to line items. At $0.36 an hour, a product team weighed whether transcription was worth the cost for every feature that touched it. At $0.10, that calculation mostly disappears: transcription becomes a default you switch on across a product rather than a cost center you audit line by line. That is a bigger change to product design than any accuracy delta Microsoft is claiming, because it removes the moment where someone has to justify using the capability at all.
VentureBeat’s framing of the swap-out strategy deserves the same scrutiny Microsoft’s benchmarks do. Microsoft has committed more than $13 billion to OpenAI and still hosts OpenAI’s models across Azure, Office, and Copilot. Replacing OpenAI’s transcription tooling with an in-house model is a narrower move than replacing OpenAI’s frontier reasoning models, and transcription is a modality where switching costs for a hyperscaler are comparatively low: audio in, text out, with no chat interface or agent framework built on top to migrate.
For any team running transcription at volume, the near-term move is a real per-language benchmark comparison rather than a decision made on Microsoft’s headline WER figure, since the 5.2% average spans 60 languages of wildly uneven difficulty. For product teams more broadly, a price this low is the signal to stop treating transcription as a feature to gate behind a paid tier and start treating it as infrastructure to run everywhere by default.
Microsoft AI’s own announcement (September 3, 2026) supplied the model’s features and self-reported benchmarks; VentureBeat’s reporting on the same date supplied the pricing comparison and the swap-out strategy analysis.