According to MarkTechPost, NVIDIA has recently introduced NemotronLabs VoiceChat 11B, an open-source, end-to-end full-duplex speech conversation model. This innovative model facilitates real-time speech comprehension and generation within a single, unified network. It obviates the necessity for a series (or chaining) of multiple models, such as Automatic Speech Recognition (ASR), Large Language Model (LLM), and Text-to-Speech (TTS), thereby substantially minimizing end-to-end latency.
