Microsoft AI has released MAI-Transcribe-2-Streaming, which turns live speech into text as audio arrives. It complements the batch model for completed recordings, producing an initial partial transcript without waiting for the recording to end. The new model is available through Voice Live API in public preview.
Final transcripts: fewer errors than Grok, but not the fastest response
At the time of publication, Artificial Analysis ranked the new model first among 38 models for final-transcript word error rate. MAI-Transcribe-2-Streaming achieved a WER of 2.5%, with final text delivered 0.13 seconds after speech ended. Grok Voice Transcribe 2.0 recorded a WER of 2.7% and latency of 0.49 seconds in the same evaluation.
WER measures errors in recognized words, with lower values indicating better results. A score of 2.5% represents roughly 2.5 substitutions, omissions or insertions per 100 reference words. Alongside this metric, Artificial Analysis tracks the time needed to produce partial and final transcripts.
The lowest error rate did not translate into the fastest final response in this evaluation. Artificial Analysis measured Cartesia Ink Preview at 0.11 seconds, compared with Microsoft's 0.13 seconds. Cartesia's WER was 3.1%, compared with MAI's 2.5%.
Artificial Analysis clocks first partial transcript at 0.12 seconds
In a separate evaluation of the first partial transcript, Artificial Analysis recorded a WER of 2.5% and latency of 0.12 seconds for MAI. Grok Voice Transcribe 2.0 registered a 3.4% error rate and a response time of 0.49 seconds in this part of the assessment.
The first partial transcript showed a similar speed-accuracy trade-off. In the same evaluation, Cartesia Ink-2 produced its first transcript in 0.07 seconds, but its WER reached 4.0%.
MAI Transcribe covers 60 languages, Microsoft says
Microsoft lists a single multilingual model supporting 60 languages for the MAI Transcribe family. Its listed features include speaker diarization and timestamps for individual words. Availability of specific features can vary by endpoint.
For recognizing names and specialized terminology, Microsoft lists an option to give specified keywords greater weight. The model family also offers a choice between verbatim transcription and cleaned-up output.
Live audio through Voice Live API, currently without an SLA
Live transcription uses the separate Voice Live API, while prerecorded audio files are processed through Fast Transcription API. MAI-Transcribe-2-Streaming is available through Voice Live API in public preview without a service-level agreement (SLA). Microsoft advises against deploying the preview in production.
The streaming model costs $0.54 per audio hour, matching Gemini 3.5 Transcribe Live and exceeding ElevenLabs Scribe v2 Realtime and Deepgram Flux, while the batch MAI-Transcribe-2 model has a promotional price of $0.10 per audio hour through the end of 2026.



