Cartesia Unveils Its Cutting-Edge, Real-Time Conversational TTS Model: Sonic-3
2025-10-29 / Read about 0 minute
Author:小编   

Cartesia has proudly declared the triumphant conclusion of a substantial $100 million funding round. This round saw participation from prominent investors such as Kleiner Perkins, Index Ventures, Lightspeed, and NVIDIA. Simultaneously, the company has officially rolled out its latest real-time conversational text-to-speech model, Sonic-3. Leveraging the innovative State Space Model (SSM) architecture, this model boasts an impressively low inference latency of merely 90 milliseconds and an end-to-end response time of just 190 milliseconds. Such performance metrics position it among the swiftest text-to-speech (TTS) systems currently on the market. Sonic-3 is versatile, supporting a total of 42 languages, and offers a wealth of expressive capabilities for paralinguistic communication. This significantly elevates the realism and immersive quality of voice interactions.