Microsoft AI Releases MAI-Transcribe-2-Streaming: The #1 Real-Time Speech-to-Text Model on Artificial Analysis
Microsoft AI has released MAI-Transcribe-2-Streaming, its first streaming speech-to-text (STT) model. It launched on October 1, 2026, alongside two text-to-speech models, MAI-Voice-2.1 and MAI-Voice-2.1-Flash. According to Artificial Analysis, the model ranks #1 out of 38 models for both final and first partial transcript accuracy. It is designed for voice agents, live captions, and dictation—applications where latency directly determines the quality of the experience.
What Microsoft Shipped
MAI-Transcribe-2-Streaming is the real-time sibling of the batch model MAI-Transcribe-2, which was released in September 2026. It transcribes 60 languages with automatic, continuous language detection. Audio streams in continuously, and text streams back while the speaker is still talking.
The model emits its first hypotheses—called partials—just over 100 ms after receiving audio. It revises those partials as additional context arrives, then commits a stable final transcript. As a result, an agent can begin reasoning or calling tools mid-sentence. Microsoft's internal tests indicate that words appear twice as fast as with its closest competitor.
What Artificial Analysis Measured
The AA-WER Streaming index uses approximately 8 hours of audio. The mix consists of AA-AgentTalk (50%), VoxPopuli (25%), and Earnings22 (25%). Latency is measured from the end of speech, as detected by SileroVAD.
(Content truncated in source; additional sections were not provided.)
via MarkTechPost
