By 2026, three prominent text-to-video companies—Zhixiang Future, Shengshu Technology, and Aishi Technology—have shifted their focus from their once-thriving AI short video services to the burgeoning world model sector, each charting a unique course for development. Zhixiang Future's HiDream-O1-Embodied places a strong emphasis on embodied perception and robust anti-disturbance capabilities, aiming to tackle interference challenges encountered in real-world environments. Shengshu Technology's Motus2, on the other hand, zeroes in on robotic manipulation, with a determined effort to boost the success rates of robotic operations. Meanwhile, Aishi Technology's PixVerse R2 stands out by highlighting real-time interactive virtual worlds, enabling seamless synchronous generation and control.
These companies bring inherent strengths to the world model domain. They can capitalize on their pre-existing visual prior knowledge, harness multimodal control capabilities, and offset the scarcity of real-world operational data through innovative data augmentation techniques. Moreover, they already boast mature underlying pipelines, providing a solid foundation for their endeavors.
However, the road ahead is not without obstacles. Video models' comprehension of the physical world is currently confined to the visual realm, creating a notable disconnect with the actual physical world. Additionally, issues such as long-term operational state drift, the cross-hardware transfer of force and tactile sensing, and the establishment of efficient data feedback pipelines remain to be addressed.
The crux of competition in the world model arena lies in system-level closed loops. Future industry dynamics are likely to revolve around collaborative efforts across various links in the chain. Furthermore, these companies also grapple with the potential transition from being mere software service providers to becoming integrated software-hardware firms, a transformation that could redefine their market position and competitive edge.
