Microsoft releases MAI-Transcribe-2-Streaming for live speech transcription

Microsoft releases MAI-Transcribe-2-Streaming for live speech transcription

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
3. 10. 2026
2 minutes reading · 13 views
Listen to the article
Audio version of the article
Microsoft releases MAI-Transcribe-2-Streaming for live speech transcription

Microsoft AI has released MAI-Transcribe-2-Streaming, which turns live speech into text as audio arrives. It complements the batch model for completed recordings, producing an initial partial transcript without waiting for the recording to end. The new model is available through Voice Live API in public preview.

Final transcripts: fewer errors than Grok, but not the fastest response

At the time of publication, Artificial Analysis ranked the new model first among 38 models for final-transcript word error rate. MAI-Transcribe-2-Streaming achieved a WER of 2.5%, with final text delivered 0.13 seconds after speech ended. Grok Voice Transcribe 2.0 recorded a WER of 2.7% and latency of 0.49 seconds in the same evaluation.

WER measures errors in recognized words, with lower values indicating better results. A score of 2.5% represents roughly 2.5 substitutions, omissions or insertions per 100 reference words. Alongside this metric, Artificial Analysis tracks the time needed to produce partial and final transcripts.

The lowest error rate did not translate into the fastest final response in this evaluation. Artificial Analysis measured Cartesia Ink Preview at 0.11 seconds, compared with Microsoft's 0.13 seconds. Cartesia's WER was 3.1%, compared with MAI's 2.5%.

Artificial Analysis clocks first partial transcript at 0.12 seconds

In a separate evaluation of the first partial transcript, Artificial Analysis recorded a WER of 2.5% and latency of 0.12 seconds for MAI. Grok Voice Transcribe 2.0 registered a 3.4% error rate and a response time of 0.49 seconds in this part of the assessment.

The first partial transcript showed a similar speed-accuracy trade-off. In the same evaluation, Cartesia Ink-2 produced its first transcript in 0.07 seconds, but its WER reached 4.0%.

MAI Transcribe covers 60 languages, Microsoft says

Microsoft lists a single multilingual model supporting 60 languages for the MAI Transcribe family. Its listed features include speaker diarization and timestamps for individual words. Availability of specific features can vary by endpoint.

For recognizing names and specialized terminology, Microsoft lists an option to give specified keywords greater weight. The model family also offers a choice between verbatim transcription and cleaned-up output.

Live audio through Voice Live API, currently without an SLA

Live transcription uses the separate Voice Live API, while prerecorded audio files are processed through Fast Transcription API. MAI-Transcribe-2-Streaming is available through Voice Live API in public preview without a service-level agreement (SLA). Microsoft advises against deploying the preview in production.

The streaming model costs $0.54 per audio hour, matching Gemini 3.5 Transcribe Live and exceeding ElevenLabs Scribe v2 Realtime and Deepgram Flux, while the batch MAI-Transcribe-2 model has a promotional price of $0.10 per audio hour through the end of 2026.

Advertisement

Content created with help from UpTier.

SEO and GEO on autopilot. UpTier’s multi-agent systems write and optimize content for search engines and AI answers.

Discover UpTier ↗

Category:AI
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

Qwen-Image-2.1 combines image generation and editing with transparent outputQwen-Image-2.1 combines image generation and editing with transparent output
Alibaba has released a single checkpoint for image generation and editing. Qwen-Image-2.1 supports transparent PNGs and, according to the company, accepts up to 10 reference images. Commercial deployment requires a separate license.
3 min read
3. 10. 2026
CoreWeave launches Forge for training and continuously improving AI agentsCoreWeave launches Forge for training and continuously improving AI agents
Forge connects model training, evaluation and improvement with insights from production. It offers post-training without a dedicated cluster, experiment analysis and isolated environments for agents.
2 min read
3. 10. 2026
Gemini 4 Argon can generate up to one million output tokens, Google saysGemini 4 Argon can generate up to one million output tokens, Google says
Google introduced Gemini 4 Argon with longer reasoning sequences. Vals says the model leads its professional-task index and uses fewer tokens than Claude Sonnet 5.5 for comparable work. Access starts with cybersecurity partners.
3 min read
3. 10. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok