Shengshu Technology has just rolled out Vidu S2, a cutting-edge video generation model tailored for real-time interaction and editing. This innovative model boasts three key capabilities: real-time digital human interaction, offline digital human video generation, and real-time video editing. Among these, S2-Avatar excels in facilitating real-time interactions with digital humans, seamlessly handling dynamic actions like dancing and even allowing the insertion of reference images mid-action for enhanced customization. Meanwhile, S2-Editing is designed to receive video streams in real-time, enabling on-the-fly alterations to style and characters while preserving the original movement rhythm. Additionally, S2-Avatar [Offline] offers the capability to generate digital human videos in batches, with movements that are fully controllable.
Vidu S2 places a strong emphasis on technological advancements in speed, stability, and controllability. It effectively tackles the challenges inherent in real-time video generation by employing techniques such as data expansion and annotation, bidirectional training with streaming conversion, Self-Replay Forcing for optimization, and a two-stage architecture that accelerates inference. Furthermore, the model introduces a VLM Agent to adeptly handle dynamic references and incorporates frame-aligned attention mechanisms to facilitate real-time video editing with precision.
In multiple rigorous evaluations, Vidu S2 has consistently showcased exceptional performance, extending its real-time generation and editing prowess to the realm of spatial videos. When compared to its predecessor, Vidu S1, Vidu S2 exhibits significant improvements in terms of control objects, resolution, and movement range. These enhancements mark a substantial leap forward in propelling AI video technology towards seamless, continuous operation and real-time responsiveness.
