TestingCatalog found a voice model called MAI Realtime tucked inside Microsoft’s MAI Playground, sitting there without any public announcement or documentation. Only a small set of partners appears able to reach it right now, based on what the outlet could see in the listing. Everything visible points to Microsoft’s first native full-duplex voice system, a category the company has not previously shipped.
Full duplex means the model keeps listening even while it talks, instead of pausing until you finish before it replies. Most voice assistants, including Copilot’s current voice mode, work in half duplex: the system goes silent while you speak, then builds a response only after you stop, which is why interrupting them produces dead air or garbled overlap. A full-duplex model runs its microphone and its speech output concurrently, so it can react mid-sentence and adjust course before you finish talking. That is a change in architecture, not a latency tweak layered onto a turn-based pipeline.
Two voices, Victoria and Grant, are available in the playground build. TestingCatalog describes both as sounding notably more human than the voice currently used in Copilot. The model supports numerous languages, including English, Spanish, French, Japanese, and Arabic, with automatic detection that can switch mid-conversation without resetting. Turn-taking runs through one of two listener modes: a Switchboard setup that relies on an MAI-Ears endpointer using inline control tokens, or a deterministic pairing of silence-based and Whisper-based semantic endpointing. TestingCatalog reports that interruptions come through cleanly and that replies stay fast. Singing and other non-speech sounds are outside its range entirely, which keeps the system focused on conversation rather than open-ended audio generation.
None of this is confirmed by Microsoft. TestingCatalog describes a playground entry a small number of partners can already test, not a public launch. Microsoft has disclosed no pricing, no release date, and no confirmation the model will ship broadly, and has not commented on MAI Realtime. The listing could still change before any wider rollout.
The model would close a real gap in Microsoft’s speech stack. MAI-Voice-2, along with its Flash variant, handles text-to-speech, and MAI-Transcribe-1.5 handles recognition, but the speech-to-speech layer inside Azure’s Voice Live API still runs on OpenAI’s GPT-Realtime model. A working MAI Realtime would let Mustafa Suleyman’s AI division swap that external dependency for its own technology, continuing a pattern that already includes seven in-house models unveiled at Build 2026 and a steady removal of OpenAI components from Copilot, Teams, and Bing. Building a native voice model instead of licensing one signals how independent Microsoft wants its AI stack to become.
Developers would likely meet this model first inside Microsoft Foundry, while Copilot’s voice mode is where consumers would encounter it, though Microsoft has attached no timeline to either. Teams evaluating voice AI vendors should treat MAI Realtime as a signal worth tracking until Microsoft confirms availability.
TestingCatalog’s Alexey Shabanov first reported on MAI Realtime on August 2, 2026.