StepFun Unveils StepAudio 3 Series of Advanced Voice Models, Achieving Multiple World-Leading Milestones
2 day ago / Read about 0 minute
Author:小编   

StepFun has proudly announced the launch of its StepAudio 3 series, a suite of five cutting-edge models: StepAudio3Realtime, StepAudio3ASR, StepAudio3TTS, StepAudio3Gen, and StepAudio3Music. These models have not only made their debut but have also soared to the top of global rankings on the Artificial Analysis leaderboard, with all now accessible via the StepFun Open Platform.
StepAudio3Realtime stands out as a pioneer, offering native full-duplex real-time dialogue capabilities. It excels in comprehending multi-dimensional audio information, executing parallel reasoning and voice generation, and performing asynchronous tool calls, securing its position as the world's best in its category on the leaderboard.
StepAudio3ASR merges high-precision speech recognition with the power of large language models, seamlessly adapting to diverse scenarios and complex inputs. With an impressively low word error rate of just 1.7%, it shares the top spot globally.
StepAudio3TTS replicates human speech patterns with remarkable accuracy, capturing paralinguistic expressions, autonomously adjusting emotional tones, and supporting streaming generation for a truly natural listening experience.
StepAudio3Gen showcases its versatility by uniformly generating complete audio sequences, encompassing both human voices and sound effects, based on textual descriptions. It offers precise control over sound characteristics and timing, ensuring a tailored audio output.
Meanwhile, StepAudio3Music revolutionizes music creation by enabling the generation of original compositions and facilitating multi-round interactive creation. It refines musical works based on various inputs, controls elements such as musical style, and models song structure for a cohesive and engaging listening experience.