2026-09-02 11:40
Introduction: Developed by Fei-Fei Li's team, achieving pixel-level camera control and sparse photo-based 3D reconstruction. What is Atlas? Atlas is World Labs, founded by Fei-Fei Li.
What is Atlas?
Atlas is the world’s first multimodal world model launched by World Labs, founded by Fei-Fei Li. The model natively understands text, images, videos, and 3D spatial information, anchoring visual content within a 3D coordinate system through spatial context. It enables pixel-level precise camera control generation, sparse photo-based 3D reconstruction, and spatiotemporal simulation. With just a few ordinary photos, the model can generate long-form videos with arbitrary camera movements and reconstruct real-world scenes, providing a simulated training environment for robotics.

Main Features of Atlas
- Camera Control Generation: Input 1–6 photos and specify camera poses to generate up to 1-minute, 1440p cinematic videos with high geometric consistency across views.
- Spatial Reconstruction: Reconstruct high-fidelity 3D point clouds or Gaussian splatting models from only 2–25 standard photos, eliminating the need for professional scanning equipment.
- Spatiotemporal Simulation: Understands spatial structure and world dynamics, supporting multi-camera “bullet time” reconstruction and enabling Real-to-Sim simulation environments for robotics training.
- Image Generation: Generate high-quality images and 360° panoramas from text or image prompts, supporting complex instructions and diverse visual styles.
Technical Principles of Atlas
- Multimodal Autoregressive Diffusion Transformer: Atlas was pretrained from scratch using a unified architecture that encodes text, images, video frames, camera poses, and 3D depth maps into a single “spatial context” sequence. Outputs are generated autoregressively, one element at a time. Simultaneously, diffusion mechanisms progressively denoise high-dimensional continuous data (images/videos), combining the sequential flexibility of LLMs with the high-fidelity rendering capabilities of video generation models.
- Spatial Context: Unlike LLMs that rely on textual context, Atlas assigns explicit 3D coordinates and depth information to each input image, constructing a 3D scene representation with spatial constraints within the model. When generating new viewpoints, the model extrapolates and completes within this established spatial framework, ensuring geometric consistency and physical plausibility across different perspectives.
How to Use Atlas
- Visit Official Website: Apply for Atlas access at https://form.typeform.com/to/zHFR4r3A.
- Generate Cinematic Videos via Camera Control: Upload 1–6 reference images, define camera pose and motion trajectory—Atlas generates up to 1-minute cinematic footage with smooth, consistent motion.
- 3D Spatial Reconstruction: Upload 2–25 ordinary photos; Atlas automatically reconstructs them into navigable 3D point clouds or Gaussian splatting scenes without requiring specialized hardware.
- “Bullet Time” Effects: Capture short clips using 3–5 regular smartphones from different angles; Atlas reconstructs the scene to enable playback from any viewpoint.
- Robotics Simulation Training: Record short videos of real environments; Atlas generates high-precision digital twins for robots to explore and test in virtual space.
- Text-to-Image / Panorama Generation: Input text prompts or reference images to directly generate high-quality images or 360° panoramas.
Core Advantages of Atlas
- Pixel-Level Camera Control: Treats camera pose as native input rather than relying on ambiguous text descriptions, enabling director-grade precision in camera movements.
- Ultra-Sparse Reconstruction: Achieves high-accuracy 3D scene reconstruction from just a few smartphone photos, far surpassing traditional scanning methods.
- Spatial Geometric Consistency: Anchors images in 3D coordinates via Spatial Context, ensuring stable, flicker-free generation across multiple viewpoints.
- Unified Multitask Architecture: A single model handles generation, reconstruction, simulation, and image synthesis—eliminating the need to switch between specialized tools.
- Real-to-Sim Closed Loop: Directly converts real-world video into robot simulation environments, drastically reducing embodied AI training costs.
- Scalability with Compute Growth: Pretrained from scratch and validated to continuously improve with increased compute, offering immense future potential.
Project Repository
- Official Website: https://www.worldlabs.ai/blog/atlas
Competitive Comparison with Similar Projects
| Comparison Dimension |
Atlas (World Labs) |
Genie 3 (Google DeepMind) |
| Positioning |
Omni world model for spatial intelligence |
General-purpose real-time interactive world model |
| Core Inputs |
Text, images, videos, camera poses, 3D depth maps |
Images, action commands, state history |
| Camera Control |
Pixel-level precision with native support for camera pose parameters |
Primarily indirect control via text/action descriptions |
| 3D Reconstruction |
Reconstructs point clouds/Gaussian splatting from 2–25 standard photos |
Focuses on interactive world generation, not specialized in sparse reconstruction |
| Video Generation |
Up to 1 minute, 1440p with exceptional spatial geometric consistency |
Emphasizes real-time interactivity, shorter segment durations |
| Physics Simulation |
Supports robot training environment generation via Real-to-Sim |
Highlights self-learning physics engines with real-time feedback |
| Output Formats |
Images, videos, point clouds, 3D Gaussian Splatting |
Interactive world states, video frames |
Application Scenarios of Atlas
- Film & VFX Production: Directors provide a few scene photos and design camera trajectories; Atlas generates high-resolution (1440p) cinematic sequences with precise camera motion, enabling "bullet time" effects using just a few smartphones instead of hundreds of professional cameras.
- Game & Virtual Environment Creation: Art teams upload a small number of real-world photos; Atlas automatically reconstructs them into navigable 3D Gaussian splatting scenes, directly importable into game engines, dramatically reducing production time for open-world or VR environments.
- Robotics & Embodied Intelligence Training: Record short videos of real environments with smartphones; Atlas generates high-precision digital twins for robots to repeatedly navigate and experiment in virtual space, addressing core challenges of slow and costly real-world data collection.
- Architecture & Real Estate Visualization: Reconstruct complete 3D models of buildings or interiors from just a few on-site photos, supporting arbitrary-angle exploration and aerial view generation for property pre-sales or heritage digitization.
- Advertising & E-commerce Content Generation: Input product images or text prompts to quickly generate 360° panoramic videos or stylized commercials; pixel-level camera control ensures consistent geometry and lighting across all camera movements.
Source: AI Tools Hub