Microsoft Releases MAI-Transcribe-2 at 10 Cents an Hour
The new model adds language coverage and transcript-formatting features while sharply lowering Microsoft’s introductory rate. Its benchmark pitch is focused on recorded-audio throughput, leaving live transcription and speaker-label accuracy as practical tests for buyers.
Listen to this story
The audio brief
Story brief
3 key pointsMicrosoft’s MAI-Transcribe-2 expands coverage to 60 languages and adds production-oriented controls while cutting the launch price to $0.10 per audio hour. The offer is substantially below Microsoft’s original $0.36 rate and is positioned for fast batch and long-form processing, with claimed speed advantages over major rivals. Buyers should treat the price as provisional: Microsoft has not disclosed the end date or...
- 01
The early-bird price is $0.10 per audio hour, versus $0.36 for Microsoft’s first transcription model.
- 02
Support grew from 43 languages in MAI-Transcribe-1.5 to 60 in MAI-Transcribe-2.
- 03
Features include diarization, word timestamps, keyword biasing, language identification, code switching, and clean or verbatim output.
Microsoft has released MAI-Transcribe-2, a speech-recognition model priced at an early-bird rate of $0.10 per audio hour. The move gives developers a cheaper Microsoft option for turning recorded speech into structured transcripts, while leaving important details about real-time use unresolved.
From April’s first model to a 10-cent launch offer
MAI-Transcribe-2 is Microsoft’s third transcription model in five months, following MAI-Transcribe-1 on April 2 and MAI-Transcribe-1.5 on June 2. Microsoft says it supports 60 languages, up from 43 for version 1.5 and 25 for the original, and is available through Microsoft Foundry and MAI Playground.
The price is the immediate change. Microsoft charged $0.36 per audio hour for its first model; MAI-Transcribe-2’s $0.10 rate is an early-bird offer. The company has not specified when that offer ends or what a standard rate will be, a consequential unknown for teams planning longer-term deployments.
Transcript controls beyond plain text
Microsoft says the model includes speaker diarization, which identifies who said what, along with word-level timestamps, keyword biasing for supplied terms and automatic language identification. It also supports configurable verbatim and clean output styles and code switching, making transcripts searchable, editable and assignable to speakers rather than leaving them as blocks of text.
Speed claims meet production limits
Microsoft says MAI-Transcribe-2 ranks first on the FLEURS benchmark across 60 languages, with a 5.2% average word error rate. The company also says it ranks second on Artificial Analysis’s word-error-rate leaderboard and sits on that site’s accuracy-latency Pareto frontier, meaning no model in the comparison is both faster and more accurate.
Artificial Analysis evaluations cited by Microsoft put the model at 10 times the speed of OpenAI’s GPT-Transcribe, seven times that of ElevenLabs’ Scribe v2, and five times that of Google’s Gemini 3.5 Transcribe. For batch transcription, that speed can increase audio-processing throughput while reducing GPU-hours for a fixed workload.
What buyers still need to test
- The release emphasizes batch throughput and long-form audio, not real-time or streaming transcription performance.
- The 5.2% word-error result averages 60 languages. Microsoft did not provide a per-language accuracy breakdown.
- Word error rate does not measure whether speakers are labeled correctly, and Microsoft did not provide a diarization error rate.
Editorial analysis
Our Read
Microsoft appears to be treating transcription as a product category where price, throughput and workflow outputs can be packaged together. The release’s strongest commercial signal is not its benchmark position alone, but the combination of a 10-cent introductory rate, 60-language coverage, diarization and timestamps. The next meaningful test is operational: Microsoft has emphasized batch and long-form audio, while recent competing coverage has highlighted live processing and speaker labeling. Watch for streaming measurements and diarization-quality figures, which would show whether the model’s low launch rate extends to those production demands.
Sources
- venturebeat.comMicrosoft AI’s MAI-Transcribe-2 undercuts OpenAI, Google and ElevenLabs on price and speed