World Labs has announced Atlas, an AI model capable of generating, reconstructing, and simulating 3D worlds from just a few photos. The company claims Atlas outperforms specialized models at their own tasks, potentially making many of them obsolete. Atlas is the first model developed by the company to achieve this at scale, focusing on spatial intelligence to understand 3D space as humans do. According to World Labs, the model processes inputs by anchoring them to specific positions in 3D space rather than treating them as flat sequences. This approach, called 'spatial context,' allows Atlas to generate new frames or viewpoints based on this shared understanding. Fei-Fei Li, co-founder of World Labs, highlighted the need for 3D-aware architectures in her November 2025 essay, arguing that current models break data into one- or two-dimensional sequences, making spatial tasks unnecessarily complex. World Labs describes Atlas as an omni-model trained on text, images, video, and 3D data, enabling it to produce new views at freely chosen camera positions and angles. The model outputs up to one minute of video at 1440p, allowing users to control every shot without relying on random outputs. Atlas also reconstructs real scenes from as few as one to several dozen input images, delivering faithful results with just two or three images and outperforming specialized 3D models. It can handle over a hundred inputs, as demonstrated in a test where it progressively assembled Stanford's Main Quad from two to 25 ground-level photos and generated aerial views above the campus. In a comparison within the OpenWorldLib framework, systems like VGGT and InfiniteVGGT showed geometric inconsistencies and blurry textures when the camera moved significantly, whereas Atlas performed more reliably. Atlas can output results as actual 3D data, including point clouds and 3D Gaussian splats, matching the representation used in Marble, the company's existing product. As a simulator, Atlas models space and time together, producing a 'bullet time' effect that allows users to view scenes from otherwise impossible angles. The model also serves as a real-to-sim tool for robotics, enabling users to simulate and vary grasping and movement tasks by swapping out objects, positions, lighting, or backgrounds. World Labs showed this approach in August 2026 with its real-to-sim-to-real engine, which created thousands of variants from a single real-world task and trained control models entirely in simulation. The technology came from SceniX, a startup World Labs acquired in July. Atlas combines ideas from both language models and video models, generating output piece by piece like a language model while using diffusion principles from image and video models to boost image quality. World Labs says no single benchmark captures what Atlas can do but points to two sets of tests, where Atlas outperformed more specialized models in camera-controlled generation and few-view 3D reconstruction. Human evaluators preferred Atlas in 75 to 94 percent of comparisons against various models, and it led in reconstruction with a median error of 25.3, ahead of Pi3X and VGGT-Ω 1B. Atlas will power future versions of Marble and other products and is currently available through an early-access program for select partners. The company says Atlas's performance improves with more training compute and expects this trend to continue as it scales. Atlas posts the lowest average reconstruction error among all compared specialized models, where lower values are better. World Labs was founded in 2024 by Fei-Fei Li, who created ImageNet and led Google Cloud's AI division from 2017 to 2018. The company was launched with funding from Andreessen Horowitz, AMD, Intel, and Nvidia. A first system was released in late 2024, though users could only move a few virtual meters before hitting invisible boundaries. Marble followed in November 2025. In February 2026, the company secured a $1 billion funding round from Autodesk, Andreessen Horowitz, Nvidia, and AMD. Bloomberg had previously reported talks at a $5 billion valuation. What counts as a world model remains contested among researchers. An international team led by Peking University proposed a unified definition in April 2026 through OpenWorldLib, excluding pure text-to-video models because they lack feedback loops with the real world. 3D reconstruction and simulators like those in Atlas qualify as core building blocks in that framework because they provide environments where physical rules can be verified.

Source: thedecoder