Addressing the limitation of current audio-video interactions being mostly confined to audio-video Q&A and lacking native audio-video conversational capabilities without the need for text or ASR transcription, relevant research has introduced the OmniVChat native audio-video conversation framework. This framework utilizes the multi-agent data engine OmniVChat-Studio to synthesize single-round and multi-round native audio-video conversation data that includes reference responses and scoring criteria. It also constructs OmniVChat-Bench, an evaluation benchmark comprising 2,800 items across five major capability categories. Meanwhile, the hierarchical scoring criteria are transformed into a reinforcement learning reward design, OmniVChat-RL, and models are trained based on this synthetic data. Experimental results demonstrate that by using only synthetic conversation data for reinforcement learning, the model's performance on untrained real human interaction datasets improves significantly, proving that carefully designed synthetic data can effectively train native audio-video conversation models for real-world scenarios.
