West Lake University Unveils the Code Video World Model: Where Code Fuels World Evolution
2 hour ago / Read about 0 minute
Author:小编   

The research team from the AGI Lab at West Lake University has introduced the innovative Code World Model, a framework that dissects world evolution and visual rendering into two interdependent yet distinct challenges. Within this model, the Coding Agent assumes the role of the 'world brain,' tasked with upholding the rules and states of the world, as well as dictating the logic of its evolution through code. This agent demonstrates remarkable adaptability, seamlessly integrating, invoking, or altering code in response to emerging scenarios. Meanwhile, the video model undertakes the responsibility of converting the evolved state of the world into visually stunning, high-fidelity scenes.

A key innovation of this model is the introduction of a Proxy, which serves as a coarse-grained visual cue to bridge the gap between world states and the video model. This Proxy strikes a delicate balance between constructibility—the ease with which it can be built or understood—and constraint strength—the degree to which it restricts or guides the model's output. To compile training data, the team synchronously records gameplay footage alongside its corresponding states, thereby generating the Proxy. Impressively, this approach is versatile enough to process real-world videos as well.

After subjecting the prototype to approximately 5.6 hours of validation using gameplay videos, the model proved capable of generating visual content that adheres to the constraints imposed by the Proxy. This enables seamless style switching during extended generation sessions, all while ensuring the continuous progression of world states. By clearly delineating responsibilities, the Code World Model effectively decouples high-level reasoning from low-level execution. This design not only renders it ideal for game generation but also positions it as a valuable environment for training and evaluating intelligent agents in domains such as embodied AI. Ultimately, this model offers fresh perspectives and insights for the development of open-world models.