NetEase Youdao Releases Open-Source Code for Two Real-Time Interactive Models, R2T2 and T3PO, from the Ziyue Live Series
1 day ago / Read about 0 minute
Author:小编   

NetEase Youdao has recently made two real-time models from its Ziyue Live series available as open-source: the Confucius4-R2T2 streaming speech recognition model and the Confucius4-T3PO streaming translation model. Both of these models have excelled, securing top positions on the ASR (Automatic Speech Recognition) and Translation Trending charts on Hugging Face. These models are designed to facilitate stable simultaneous speech recognition and translation. Specifically, the R2T2 model operates with an average latency ranging from 200 to 600 milliseconds, ensuring swift and responsive speech recognition. Meanwhile, the T3PO model prioritizes translation quality while also maintaining low latency, providing accurate and timely translations. Since their release as open-source, these two models have garnered adaptation support from various platforms and tools, including ZeroGPU Demo, audio.cpp, GGUF quantization, and llama.cpp/Ollama, among others, further enhancing their versatility and applicability.

  • C114 Communication Network
  • Communication Home