Microsoft AI released MAI-Transcribe-2 on 3 September and priced it at $0.10 per hour of audio through 31 December 2026. That works out to about $1.67 per thousand minutes. The company has not published the rate that applies on 1 January.
The accuracy claim is the part that makes the price interesting. Microsoft says the model ranks first on the FLEURS benchmark across 60 languages with an average word error rate of 5.2%, ahead of Gemini 3.5 Transcribe, GPT-Transcribe, Whisper V3-Large and ElevenLabs' Scribe v2. On speed it claims up to 10x faster than GPT-Transcribe, 7x faster than Scribe v2 and 5x faster than Gemini 3. It ships with speaker diarization, word-level timestamps and configurable transcription styles.
These are vendor benchmarks on a vendor-chosen suite, and we have spent this week making exactly that point about a much larger model. Treat 5.2% as a claim, not a measurement. But the price is not a claim. The price is a number you can put in a spreadsheet today.
The floor moved, and floors do not move back quietly
Speech-to-text has been the least differentiated layer of the AI stack for two years. Everyone's model is good enough for meeting notes; nobody's is good enough for uncorrected legal transcription. When a category is undifferentiated on quality, it competes on price, and the first vendor to make price the headline is the vendor that has decided margin is not where it plans to win.
Microsoft does not need transcription revenue. It needs audio flowing through Azure, where the compute, the storage and the downstream summarisation all bill separately. Ten cents an hour is a loss leader for a pipeline, and the promotional end-date is the tell: this is a customer-acquisition price with an expiry, not a cost structure.
Our take: if transcription is a line item in your product's COGS, your unit economics changed on Wednesday and you did not do anything to earn it. Reprice your own offering deliberately rather than pocketing the spread and forgetting about it, because your competitor is reading the same announcement. The trap is the date. Anything you architect around $0.10 needs to survive whatever the January number is, and you have no visibility into that number. Build the migration path before you build the dependency.
The preview label matters
MAI-Transcribe-2 is in public preview in Azure Speech, without a service-level agreement and explicitly not recommended for production workloads. That is not boilerplate you can wave through. A preview service with no SLA, at a promotional price, with an undisclosed post-promotional rate, is three separate pieces of uncertainty stacked on the same dependency — and Amazon has just reminded everyone what happens when a hosted model reaches end of life and the platform does not migrate you.
The sensible pattern is the one that has worked all year: route through an abstraction, benchmark on your own audio rather than on FLEURS, and keep a second provider warm. Your audio is not the benchmark's audio. Accented speech, overlapping speakers, bad microphones and domain jargon are where word error rates go from 5% to 20%, and none of that is in an average across 60 languages.
What to watch
- The January price. The single most important unpublished number in this launch.
- Whether OpenAI, Google or ElevenLabs match. A response within weeks confirms a price war; silence suggests they think the accuracy claim will not hold on real audio.
- Independent benchmarks. Third-party speech-to-text evaluations on messy production audio, not curated multilingual sets.
- GA and an SLA. Until both exist, this is a pilot budget, not a platform decision.
Google is shipping a new Flash model every few weeks. The cadence and the pricing are the same strategy from two directions: make the commodity layer free enough that nobody shops, and sell everything sitting on top of it.
