On September 3, Microsoft launched its cutting-edge AI speech-to-text model, MAI-Transcribe-2. The model is competitively priced at $0.10 per hour, a rate that will remain in effect through the end of 2026. Designed to support 60 languages, MAI-Transcribe-2 adeptly handles conversations that involve natural code-switching—a feature where speakers effortlessly switch between languages mid-conversation.
When compared to other mainstream models, MAI-Transcribe-2 stands out for its superior accuracy, speed, and cost-effectiveness. On the FLEURS test set, it achieves an impressive average word error rate of just 5.2%, substantially outperforming its counterparts. Notably, it surpasses models like GPT-Transcribe in terms of processing speed, making it a top choice for developers seeking efficiency.
MAI-Transcribe-2 also introduces a range of innovative features. These include speaker separation, which distinguishes between different speakers in a conversation; word-level timestamps, which provide precise timing for each spoken word; and support for configurable transcription styles, allowing users to tailor the output to their specific needs. Additionally, the model offers keyword biasing, which prioritizes certain words during transcription, and automatic language identification, which seamlessly detects the language being spoken.
Currently, developers have the opportunity to access or trial MAI-Transcribe-2 on multiple platforms, enabling them to experience its advanced capabilities firsthand.
