The world’s first multimodal world model is here! “AI godmother” Fei-Fei Li marks a new milestone!

Under “AI godmother” Fei-Fei Li, her startup World Labs has released the world’s first multimodal world model—Atlas. Rather than just generating images and text, this model can produce pixel-precise images and video frames, and reconstruct them into a 3D world.

Atlas’s core capabilities revolve around “understanding and simulating 3D worlds,” aiming to help AI truly comprehend the physical laws and geometric structure of 3D space—not merely generate 2D visuals that look plausible. The underlying innovation lies in replacing “text context” with “spatial context.” Unlike large models that process text context, Atlas uses a multimodal autoregressive diffusion Transformer architecture to anchor each input image and video to an exact location within a 3D space, along with the corresponding camera poses and depth information. In effect, it gives the model a 3D sketchbook with “coordinate axes” and a “scale,” providing clear geometric and physical references for how the AI understands the world.

Atlas’s deeper significance lies in bridging the path toward “embodied intelligence.” Li Fei-Fei herself regards the model as a “milestone” achievement for her team: “So far, it is the best camera-controlling world model, opening the door to many application scenarios—from visual effects to robotics.” It’s easy to foresee that AI that understands the real world from a 3D perspective—like humans—can perceive time and space, which will accelerate the deployment of multiple physical AI scenarios.

It’s worth noting that although Atlas demonstrates strong capabilities, it still has clear technical boundaries and is not a mature, general-purpose world simulator. For example, it is good at constructing geometrically coherent static spaces, but its ability to perform general physics simulations—such as object collisions and complex dynamics—still needs validation. When the camera turns to areas that were never filmed, the model still needs to “imagine,” relying on prior knowledge to fill in missing content. The generated results may not be an exact replication of the real world, and as the generation path lengthens, there may also be local geometric drift or texture flickering. At present, Atlas is in early access and is only available to certain partners; in the future, it will be used to power products such as World Labs’ Marble.