On the a16z podcast, the co-founders of World Labs—Fei-Fei Li, Justin Johnson, Ben Mildenhall, and Martin Casado—introduced their latest innovation: Atlas. Built around the concept of “novel view synthesis,” Atlas merges generative capabilities with 3D reconstruction, allowing it to craft and reconstruct 3D scenes using just a few images or textual descriptions. This technology can even replicate the iconic “bullet time” effect with as few as three smartphones.
Unlike conventional models, Atlas is inherently multimodal, accepting various types of input and offering precise control over spatial elements. The team detailed their journey from the earlier Marble model to Atlas, highlighting that computational power remains a key obstacle in scaling up the model. Atlas promises to boost efficiency in creative and design industries and can convert real-world settings into virtual environments for robotics, helping to overcome data-related challenges.
Currently, Atlas has basic dynamic capabilities, with plans to enhance dynamic simulation, editability, and interactivity in the future. The team envisions “novel view synthesis” as a fundamental step toward artificial intelligence, on par with the significance of “next token prediction” in large language models.
