Atlas is a unified spatial intelligence model that reconstructs real scenes from 2–3 images, generates 1440p video along camera paths, and supports real-to-sim robotics workflows using Gaussian Splats to create photorealistic 3D digital twins from phone scans. This represents a significant step toward scalable spatial understanding for physical AI and robot training.
Stay in the loop on research in AI and physical intelligence.
Atlas Turns Phone Scans Into Robot Worlds—But Not Yet Into Physics.
Atlas: A World Model for Spatial Intelligence • World Labs • September 1, 2026.
The announcement behind the demo.
World Labs has introduced Atlas, a multimodal model that combines camera-controlled generation, sparse-view reconstruction, video reframing, depth prediction, and explicit 3D output. Give it one or more reference images, and Atlas can generate views along a specified camera trajectory. World Labs shows videos lasting up to one minute at 1440p, reconstructions made from a handful of photographs, and scenes exported as point clouds or 3D Gaussian splats. Atlas is currently in early access with selected partners rather than generally available.
For roboticists, the important demo is lower down the announcement. World Labs recorded two large environments with ordinary phone video, selected 24 frames from each recording, reconstructed the spaces, and then generated RGB and depth observations from the viewpoints of simulated body-mounted robot cameras. The camera could follow different navigation trajectories while the reconstructed environment remained visually consistent.
That distinction is worth preserving. World Labs says Atlas can often reconstruct a scene faithfully from two or three images, but its showcased robotics environments used 24 frames each. The announcement is not demonstrating an entire manipulation-ready simulator recovered from two snapshots. It is demonstrating a potentially much cheaper route from casual capture to a spatially coherent visual environment.
The headline, then, is not that Atlas has solved robot simulation. It is that the expensive front half of the simulation pipeline—capturing a site, reconstructing its appearance, and rendering robot-specific sensor views—may be becoming generative.
One model, one spatial context.
World Labs describes Atlas as a multimodal autoregressive diffusion transformer pretrained from scratch. Its supported inputs currently include text, RGB images, camera poses, and depth maps; video is treated as a sequence of images. Every visual observation can be attached to an explicit camera position and orientation, producing what the company calls a spatial context.
Here is the simplest mental model. Atlas receives a history of observations placed at known positions in space. You then request another observation at a target camera pose. Across the sequence, the transformer predicts one multimodal element after another. Within each generated image or depth map, a rectified-flow diffusion process turns noise into the requested output.
That is different from asking a conventional video model to “pan left.” A language instruction describes the desired motion approximately. Atlas receives the camera geometry itself. This gives the model a coordinate system around which reconstruction and generation can be organized.
It also helps explain the word unified. Sparse reconstruction, camera-controlled video, depth prediction, and novel-view synthesis become different arrangements of the same basic sequence: observations, their poses, a requested pose, and the desired output modality. AI researcher Bingyuan Liu’s technical reading of Atlas emphasizes precisely this outer-autoregressive, inner-diffusion structure, while also noting that details such as attention patterns, camera encoding, context selection, and the geometry conversion stage remain undisclosed.
Crucially, Atlas does not always construct a Gaussian splat before generating an image. World Labs co-founder Justin Johnson clarified that many results are generated directly as frames; explicit 3D is an optional output when a downstream workflow requires it. The spatial memory can therefore live implicitly in the posed context and model activations, rather than exclusively in a persistent scene representation.
Reconstruction is not the same as measurement.
Atlas sits on an interesting boundary between reconstruction and generation. If the input photographs observe a surface, the model can infer its depth and preserve its appearance. If the requested viewpoint exposes an unseen region, Atlas uses its learned prior to invent a plausible completion.
World Labs presents this as a controllable continuum. Additional images constrain the model and leave less for it to imagine. Atlas can also operate in a more conservative mode that predicts depth only for visible input pixels, leaving occluded areas unknown rather than filling them in.
This is powerful for content creation, but it complicates the term digital twin. A generated wall behind the camera may be spatially plausible without matching the real building. An inferred tabletop may look level while being wrong by several centimeters. For cinematography, that error may be invisible. For grasp planning, clearance checking, or safety validation, it can be decisive.
The public material does not establish whether Atlas’s completed geometry is metrically accurate, how uncertainty is represented, or whether generated regions can be cleanly distinguished from directly observed regions in exported assets. Those are not secondary product details. They determine whether Atlas is producing a photorealistic spatial hypothesis or an engineering-grade twin.
Why Gaussian splats matter—and where they stop.
A 3D Gaussian splat represents a scene with many colored, translucent ellipsoids distributed through space. Their positions, sizes, orientations, opacity, and view-dependent appearance are adjusted so that, when projected into a camera, they reproduce the captured photographs. The representation became popular because it can preserve difficult visual details while supporting real-time novel-view rendering.
That makes splats excellent for the observation side of robotics simulation. A wrist camera can move through a captured lab and see realistic reflections, clutter, lighting, and texture. A navigation policy can encounter visual statistics much closer to its deployment site than those produced by a manually modeled game environment.
But Gaussian splats are not automatically collision geometry. They do not inherently provide object identities, joints, mass, friction, contact surfaces, or deformation parameters.
World Labs’ earlier Marble integration with NVIDIA Isaac Sim makes this separation explicit. The workflow exports a PLY Gaussian splat for visual rendering and a separate GLB triangle mesh for collisions. The splat makes the kitchen look real; the collider stops the robot from driving through the counter. Physics, scale alignment, lighting, robot assets, and controllers are then configured inside Isaac Sim.
Atlas may automate much more of the capture and geometry-generation process, but photorealistic rendering should not be confused with physically faithful simulation.
The robotics story is a hybrid stack.
The broader context is World Labs’ July 2026 acquisition of SceniX, followed by demonstrations of a real-to-sim-to-real engine for robot training and evaluation. World Labs reported simulations spanning rigid, articulated, and deformable objects, along with policies trained entirely in simulation that transferred to several physical robot platforms. In one evaluation, each policy checkpoint received 2,000 simulated trials and 100 hardware trials, with simulation preserving the relative ranking of checkpoints.
Those results are adjacent to Atlas, but they should not be attributed to Atlas alone. The Atlas announcement does not disclose how the new foundation model connects to the SceniX physics, system-identification, policy-training, and evaluation stack.
Recent work from SceniX and Columbia University illustrates the likely hybrid direction. That system uses Gaussian splatting to render photorealistic robot observations, while a separate PhysTwin-based spring-mass model and custom physics engine reproduce deformable-object dynamics. Phone scans provide appearance; interaction videos help identify physical behavior; alignment procedures connect reconstructed splats to robot and object coordinate frames.
This division of labor is technically sensible. Atlas can supply spatial priors, depth, scene completion, novel views, and visual variation. Classical or learned simulators can handle state transitions, contact, forces, and constraints. System identification can tune the latter against real interaction data.
World Labs itself places Atlas between a renderer and a simulator. The model’s published interface does not accept robot actions as a native modality, and it does not directly output robot actions as a planner would. It can help construct worlds in which another policy learns, but it is not yet the closed-loop robot brain deciding what to do next.
What is real, and what remains missing.
The strongest part of Atlas is the native treatment of camera geometry. Exact camera poses provide a much better control interface than cinematic language, and the ability to use the same model for RGB, depth, reconstruction, and generation could remove several brittle handoffs from current pipelines.
The sparse-capture results are also consequential. Even when they require human validation and cleanup, reducing environment authoring from a specialist scanning project to a phone recording changes which sites can economically be simulated.
The benchmark evidence is less conclusive. For camera control, Atlas receives the actual geometric trajectory while competing video models receive text descriptions of the move. That is a reasonable product-level comparison—Atlas really does have a better control interface—but it is not an apples-to-apples test of underlying model capacity. World Labs’ reconstruction benchmarks are also company-run and have not yet been accompanied by public evaluation code or an independently reproducible Atlas implementation.
Temporal behavior remains another frontier. In launch discussion, Johnson acknowledged that many examples effectively move the camera through frozen time. Atlas can represent some motion, such as traffic or waves, but the team considers temporal modeling an area for further improvement. That matters because persistent state and action-conditioned change are prerequisites for a general simulator.
As of September 11, 2026, World Labs has not disclosed Atlas’s parameter count, context length, training-data composition, inference cost, detailed scaling curves, or complete architecture. Latency varies with resolution, context size, hardware, and diffusion steps, and the company envisions a faster iterative mode followed by a longer production-quality generation pass.
What Atlas changes.
Atlas is most immediately credible as a world-ingestion and sensor-generation layer.
For navigation and perception, that may be enough to matter quickly. Teams could scan deployment sites, reconstruct their visual distribution, vary lighting and clutter, and generate observations along trajectories that were never physically recorded.
For manipulation, Atlas solves only part of the problem. Contact-rich tasks still need object segmentation, editable geometry, articulation, calibrated scale, physical parameters, collision models, and action-conditioned state transitions. The World Labs–SceniX combination is interesting precisely because it can pair a generative spatial model with those more structured components.
The key metric will not be reconstruction beauty. It will be whether policies trained or selected in Atlas-assisted worlds perform better on hardware, with fewer physical trials and fewer undetected failure modes.
What should we watch next? Metric-geometry evaluations. Uncertainty masks for hallucinated regions. Automatically generated colliders and scene graphs. Persistent object state. Action-conditioned rollouts. Independent sim-to-real studies. And, above all, evidence that the system preserves the failure boundaries that roboticists actually care about.
Atlas does not yet replace Isaac Sim, MuJoCo, a system-identification pipeline, or a robot policy. What it may replace is a large fraction of the manual work required before any of those tools become useful.
That is still a significant shift: Atlas changes the economics of building robot worlds before it changes the physics inside them.