World Labs has unveiled Atlas, a multimodal AI model designed to generate, reconstruct, and simulate scenes in three dimensions. The World Labs Team announces Atlas in an official blog post, describing it as a system that works with text, images, video, camera positions, and depth data in a shared spatial context.
The company says Atlas can create new image and video views from one or more reference images while following exact camera instructions. Rather than relying only on written prompts such as “pan left” or “fly overhead,” users can provide precise camera geometry, including position, angle, and movement path. World Labs says this enables pixel-level camera control and video output of up to one minute at 1440p resolution.
For creative teams, the main appeal is the ability to direct a scene more deliberately. A user could begin with a still image, define a camera route, and generate a sequence that moves through the resulting environment. Atlas is also intended to fill in areas not visible in the original material. In a single photo of a poolside scene, for example, the model may infer the reverse side of objects or extend the surrounding landscape.
One model for generation and reconstruction
Atlas is built to handle both invented and real environments. It can reconstruct a space from a small number of images, then generate new viewpoints from within that space. World Labs says additional source images reduce the amount of visual information the model must infer. The company claims that two or three images can often produce a faithful reconstruction, while larger sets of more than 100 images can provide more detailed coverage of real locations.
The model can also produce explicit 3D assets rather than only flat images or videos. According to World Labs, its outputs include point clouds and 3D Gaussian splats, a format that represents scenes as many small, renderable elements. These assets can be useful in workflows for visual effects, game development, design, and robotics, where teams need a navigable digital representation of a location.
World Labs positions Atlas as a “world model,” meaning a system meant to understand both a scene’s layout and changes over time. One proposed use is video reframing. The company says footage from three to five ordinary cameras can be used to reconstruct an event and create new angles, including frozen-time views similar to bullet-time effects.
Another target is robotics. Atlas can reconstruct environments from phone video, then generate the RGB images and depth information that a simulated robot would see along a planned route, World Labs says. The company also says the system can support simulations of manipulation tasks by varying objects, lighting, backgrounds, and robot movements.
Technical claims and access
Technically, Atlas combines several established AI approaches. It is an autoregressive transformer, which generates one element of a sequence after another, and a diffusion model, which creates images by progressively removing noise. Images and depth maps are linked to camera poses, allowing the model to treat spatial positioning as a core input rather than an added instruction.
World Labs reports that Atlas outperforms specialized models in internal evaluations of camera-conditioned generation and sparse-view 3D reconstruction. For camera tests, the company says third-party human raters compared how well models followed requested cinematic movements. The post does not provide the full benchmark tables or independent replication.
Atlas is entering early access with selected partners. World Labs says it will use the model in future versions of Marble, its 3D world-generation product, and in other company offerings.
Stay up to date
AI for content creation: the latest tools, tips and trends. Every two weeks in your inbox: