Microsoft AI released its first streaming transcription model on October 1, a system that turns speech into text while the person is still talking, and says it now sits at the top of the Artificial Analysis accuracy ranking. The same announcement added two text-to-speech models, giving the company a complete hear-and-reply kit for voice agents.

The transcriber is called MAI-Transcribe-2-Streaming. According to Microsoft AI’s own announcement, it covers 60 languages and detects a switch between them on its own, without being told. Microsoft AI says it ranks first on Artificial Analysis, an outside leaderboard, for both finished and in-progress transcripts. It also claims a spot on the Pareto frontier, which simply means no rival on that chart is both more accurate and faster.

The speed trick is the real product. The model emits rough guesses, which Microsoft calls “partials,” roughly 100 milliseconds after audio arrives, then corrects them as more of the sentence lands. A voice agent can therefore begin reasoning or calling a tool before the caller stops speaking. For dictation and live subtitles, Microsoft AI says its internal tests show words appearing twice as fast as with its closest competitor. That comparison is the company’s own, and the post does not name the rival.

Pricing is where this gets practical. Microsoft AI charges $0.54 for every hour of audio the transcriber handles, an introductory rate that lapses when 2026 ends. The post gives no price for after that, so any team budgeting for 2027 is planning around a number that does not exist yet.

The two voice models cover the other half of the loop. MAI-Voice-2.1 handles 23 languages and 26 regional variants, and Microsoft AI says a single synthetic voice can move between them, English to Mandarin to German, while taking on a native accent in each. It costs $22 per million characters. A tutoring app could keep one teacher’s voice across a whole lesson, even as the lesson changes language.

MAI-Voice-2.1-Flash is the cheaper, faster sibling at $15 per million characters. Microsoft AI says it produces 45 seconds of audio with 150 milliseconds of end-to-end delay, runs inference 55 percent faster, and costs about 60 percent less than comparable models. Those comparisons are again Microsoft’s own, and the post does not say which models it measured against.

Both voice models can copy a speaker’s voice from a short sample, only a few seconds long, and reuse it in any supported language. Microsoft AI says consent guardrails are built in to prevent misuse, but the announcement does not describe how they work. Cloning a voice from so little material is exactly the capability fraud teams worry about, so the mechanism matters more than the promise.

Microsoft AI frames the whole package around a time budget. A voice agent must hear, understand, decide and speak inside the window where a human still feels the exchange is a conversation. Every millisecond saved on transcription and speech output is a millisecond the agent can spend thinking or checking its answer.

The competitive angle is plain: Microsoft is selling the pipes of voice automation, not a finished call-center product. The company also built a demo called Chatter in its MAI Playground to show the two models working together, though the post does not report how that demo performs under load.

One caveat on the ranking. Artificial Analysis is an independent listing, which makes the first-place claim stronger than most launch-day numbers, but the post gives no score, no margin and no list of the models it beat. The latency and cost advantages come from Microsoft AI’s internal evaluations alone.

Teams running call-center or tutoring agents should test the transcriber on their own noisy audio against their current vendor before the introductory price expires, because a 100-millisecond head start is only worth paying for if it survives accents and bad phone lines.

Reported by Microsoft AI on 1 October 2026.