On October 2, 2026, Microsoft made a significant announcement, introducing three cutting-edge models: the speech-to-text marvel MAI-Transcribe-2-Streaming, alongside the text-to-speech innovators MAI-Voice-2.1 and MAI-Voice-2.1-Flash. MAI-Transcribe-2-Streaming boasts support for 60 languages, offering seamless low-latency real-time transcription with automatic language detection capabilities. This model achieves remarkable performance metrics, including a final word error rate of 2.5%, a first-partial transcription error rate of 2.8%, and a swift final transcription time of just 0.13 seconds. Priced at an accessible $0.54 per hour of audio (with a promotional rate extending until the end of the year), it presents an attractive option for users. Meanwhile, MAI-Voice-2.1 and MAI-Voice-2.1-Flash are competitively priced at $22 and $15 per 1 million characters, respectively. Microsoft envisions developers harnessing the power of these models to construct voice assistants and conversational agents that can effortlessly listen, process, and respond in real-time with natural-sounding speech.
