Li Feifei’s Team Unveils Next-Gen World Model Atlas: Trained from the Ground Up, It Seamlessly Integrates Video Generation, 3D Reconstruction, and Robotic Simulation
2 day ago / Read about 0 minute
Author:小编   

World Labs, the startup founded by renowned AI expert Li Feifei, has introduced Atlas, a groundbreaking new-generation world model touted as the "world’s first multimodal world model." Built entirely from scratch, Atlas demonstrates a unified understanding of text, images, videos, and 3D data, situating them within a cohesive three-dimensional framework. Its impressive capabilities include generating novel viewpoints from static photos, producing videos up to one minute long in 1440p resolution along predetermined camera paths, transforming photos and videos into point clouds and 3D Gaussian splatting scenes, and much more. Furthermore, Atlas supports multi-angle visualization of dynamic events captured using a standard smartphone and can be deployed in robotic Real-to-Sim applications.

The model leverages a multimodal autoregressive diffusion Transformer architecture, with camera poses incorporated as native inputs. This design enables Atlas to achieve finer camera control precision compared to some competing models. Currently, Atlas is undergoing early-stage testing with a select group of partners and is set to serve as the foundational model for Marble, an upcoming 3D world generation platform. However, key details such as the model’s parameter scale remain undisclosed, and it has yet to undergo independent verification, leaving some room for improvement before it can fully replicate real-world dynamics.

Previously, world models such as Genie 3, Marble, and Cosmos 3 have already made their mark. What sets Atlas apart is its ability to consolidate generation, reconstruction, and simulation tasks into a single foundational model centered around 3D space, bringing us one step closer to realizing a unified world model that integrates renderers, simulators, and planners.