On September 16, iFLYTEK made a significant announcement by introducing Spark-Audio-1.0-Preview, a cutting-edge speech foundation model meticulously trained with computing resources sourced entirely from domestic suppliers. This versatile model is designed to seamlessly accept both textual and vocal inputs, while delivering outputs in text format. It boasts a wide array of capabilities, including precise speech-to-text transcription, recognition of multiple languages and dialects, identification of environmental sounds and speakers, sentiment analysis, and the ability to provide detailed responses to complex audio queries. With support for an impressive 99 languages and 202 dialects, Spark-Audio-1.0-Preview marks a pivotal step in the advancement of intelligent speech technology, transitioning from mere 'clear hearing' to 'accurate understanding'. Presently, the model is accessible for trial purposes, with plans to subsequently roll out its API on the iFLYTEK Open Platform.
