World Labs unveils Atlas, a single AI model that generates, reconstructs, and simulates 3D worlds from just a few photos
figure — world Labs, co-founded by AI researcher Fei-Fei Li, has announced Atlas, a world model that generates, reconstructs, and simulates 3D scenes from just a few images. The company claims it beats specialized models at their own tasks, which could make many of them unnecessary.
Since its founding, World Labs has pursued the goal of “spatial intelligence”, the idea that AI should understand 3D space the way humans do. Atlas is the company’s first model built to do that at scale. Rather than producing flat images or video clips, it grasps how a scene looks from any angle and how it changes over time. World Labs describes Atlas as an omni-model trained from scratch on text, images, video, and 3D data. Every input gets anchored to a specific position in 3D space rather than processed as a flat sequence. The company calls this shared spatial understanding “spatial context,” and it’s what the model uses to generate each new frame or viewpoint. According to World Labs, this anchoring separates Atlas from pure language or video models. Fei-Fei Li laid out this exact problem in a November 2025 essay. Current multimodal language models and video diffusion models break data into one- or two-dimensional sequences, she argued, which makes even simple spatial tasks needlessly hard. What’s needed are architectures that organize tokenization, context, and memory in a 3D- or 4D-aware way. For camera-controlled generation, Atlas takes one or more images and produces new views at freely chosen camera positions and angles. Camera movement is passed as a direct geometric input rather than described through text prompts, as many video models require. The model outputs up to one minute of video at 1440p. Users can control every shot themselves instead of “pulling the lever on a slot machine,” as World Labs put it, drawing a line between controlled generation and random output. For spatial reconstruction, Atlas rebuilds real scenes from as few as one to several dozen input images without special capture equipment.