Microsoft’s first native real-time voice model appears to exist — and it showed up not in a keynote but as a hidden entry in the company’s own MAI Playground. TestingCatalog reported Sunday that a listing for “MAI Realtime” is visible to a small group of partners with hands-on access. Two voices so far, Victoria and Grant. Seventeen languages, with automatic detection and mid-conversation switching. Microsoft has announced nothing: no benchmarks, no pricing, no timeline, no confirmation the thing ships at all.
Full duplex is the technically interesting part. Every voice assistant most people have used trades turns like a walkie-talkie — you stop, it starts. A bidirectional model listens and speaks simultaneously, which is the difference between an interruption being handled and an interruption being a collision. The reported build exposes two turn-taking configurations: a “Switchboard” mode built on an MAI-Ears endpointer driven by inline control tokens, and a deterministic setup pairing silence-based endpointing with a Whisper semantic endpointer. A debug panel surfaces live latency and processing steps. It does not sing or produce non-speech audio — this is a conversation engine, not a general audio model.
The strategic part is what it replaces. Microsoft already builds both ends of its voice stack in-house: MAI-Voice-2 and its Flash variant handle synthesis, MAI-Transcribe-1.5 handles recognition. The middle — the speech-to-speech layer inside Azure Speech’s Voice Live API — still runs on OpenAI’s GPT-Realtime. Microsoft owns the ears. Microsoft owns the mouth. Microsoft rents the part where they meet. MAI Realtime is shaped exactly like the thing that stops the renting.
Our take: The story here isn’t voice quality, it’s a procurement line. Mustafa Suleyman’s team has been pulling OpenAI components out of Copilot, Teams and Bing one at a time, and full-duplex speech was the conspicuous hole in an otherwise vertically integrated stack — the same pattern we flagged when Microsoft built its Mythos answer out of its rivals’ models. Caveat it properly: this is an unlisted test build, and hidden playground entries have died quietly before. But the direction of travel is not ambiguous. Every dependency Microsoft closes is a line item OpenAI stops billing — and Microsoft remains simultaneously OpenAI’s largest partner and its most motivated replacement.
What to watch
- Whether it ships, and where first. Microsoft Foundry is the obvious developer destination and Copilot voice the obvious consumer surface. Foundry availability would mean enterprise voice agents priced by Microsoft rather than resold from OpenAI.
- Latency, published. Full duplex is a milliseconds business — the debug panel exists because latency is the product. No number has been published, and until one is, this is a demo.
- Per-minute pricing. Voice agents get bought by the minute, and the floor is already moving after xAI repriced the call center at roughly $3 an hour. A first-party Microsoft model competing on cost pushes it lower again.
- The pattern, not the product. Model prices keep falling everywhere — OpenAI cut its cheapest tier 80% last month. Vertical integration is how a platform stops paying someone else’s margin on a commodity.
