Atlas: World Labs' new omni world model for spatial intelligence
World Labs' Atlas is a multimodal world model enabling spatial reconstruction, generation, and simulation for robotics and beyond.
2 min read
World Labs introduced Atlas on September 1, 2026. Atlas operates on text, images, video, camera poses, and 3D depth maps, grounding all inputs in a shared spatial context to reconstruct, generate, and simulate spatial environments.
What Atlas actually does
Atlas is a multimodal autoregressive diffusion transformer that combines multimodal inputs into a shared spatial context. It can reconstruct real-world scenes from as few as two or three input images, while also supporting one to dozens of input images. The model supports pixel-perfect camera control, producing images and videos with precise camera control. It can generate videos of up to one minute at 1440p resolution and create 360° panoramas from text or image prompts.
Atlas also models space and time, supporting Real-to-Sim workflows for robotics by reconstructing environments and generating RGB/depth observations from simulated robot viewpoints. It can aid manipulation simulation by helping create varied virtual environments for robotics training and testing, including variations in objects, positions, robot motion, lighting, and backgrounds. Atlas can also produce explicit 3D outputs such as point clouds and 3D Gaussian splats.
Why spatial intelligence matters
For robots to operate effectively in dynamic, unstructured environments, they need more than obstacle detection. They need to understand space in a way that supports simulation, training, and evaluation. Atlas provides a foundation for these capabilities by enabling realistic reconstructions and simulations of physical spaces.
How it compares to existing systems
SLAM systems primarily estimate a robot’s pose while constructing a geometric map. Atlas, in contrast, is designed as a general multimodal world model that can reconstruct, generate, and simulate spatial environments. Atlas is not a direct replacement for SLAM; it targets broader generation, reconstruction and simulation capabilities.
Broad scene coverage: Atlas is designed to operate across varied environments and scene types.
Camera-controlled generation: Produces images and videos with precise camera control, including videos up to one minute at 1440p.
Spatial reconstruction: Can reconstruct real-world scenes from one to dozens of input images, typically producing faithful reconstructions with two or three images.
Space-time simulation: Models space and time from video, supporting Real-to-Sim workflows for robotics.
Text-to-image and image-to-image generation: It can generate images and 360° panoramas from text prompts.
Limitations and open questions
Atlas is currently entering early access with select partners. Sparse-view reconstruction can imagine unseen regions, meaning the generated result isn’t necessarily an exact reconstruction of reality. More input images provide more context and reduce the amount the model has to infer. The benchmark results are reported by World Labs; for camera-controlled generation, third-party human raters were used to judge which model better followed the intended camera path.
What comes next
Potential future directions could include more efficient inference, tighter integration with reinforcement learning, or multi-agent spatial coordination, although these are not announced Atlas features. As world models like Atlas mature, we may see robots that don’t just navigate spaces but use simulated environments to support planning and training. That could redefine how machines interact with the physical world, moving beyond static maps to dynamic, generative understanding.
Source: World Labs, Atlas: A World Model for Spatial Intelligence, September 1, 2026.
Building something with AI? Let's talk.
I design and ship production AI and full-stack products for US teams. See how I can help.
View all servicesJoin the newsletter
Be the first to read our articles.

