SceneMosaic: Accelerating and Dynamizing 3D Scene Generation for Diverse Layout Simulations from a Single Image
15 hour ago / Read about 0 minute
Author:小编   

To tackle the key challenges in constructing simulation environments for robotic applications in household settings—specifically, the sluggish pace of text-to-3D generation for agents, the limited physical fidelity of parameterized image-to-3D conversion methods, and the struggles of both methods in creating varied layouts that adapt to the dynamic nature of real-life human environments—the team spearheaded by Bo Dai from the University of Hong Kong has introduced the SceneMosaic framework. This innovative approach perceives scenes as mosaics constructed from independent local patches, swiftly generating a plethora of physically compliant, semantically plausible, and functionally coherent simulation scenes through a three-stage process: structured scene reconstruction, agent layout evolution, and diverse scene composition.

Experimental outcomes reveal that SceneMosaic achieves semantic quality on par with the leading baselines while substantially minimizing physical inconsistencies. Crafting a single variant scene now takes a mere 0.03 hours, marking a significant acceleration—several dozen times faster—compared to existing methods, and its diversity performance far surpasses that of traditional perturbation sampling techniques. At present, SceneMosaic is limited to generating single-room-scale environments but is poised to expand its capabilities to encompass residential-scale, large-scale simulation environment construction in the near future. The project is now available as open-source software.