New Atlas model helps train robots and create effects in 1440p

New Atlas model helps train robots and create effects in 1440p

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
3. 9. 2026
5 minutes reading · 1 views
Listen to the article
Audio version of the article
New Atlas model helps train robots and create effects in 1440p

World Labs, led by Stanford professor Fei-Fei Li, has unveiled its new world model, Atlas. It can create a photorealistic three-dimensional environment from text, images, and video, allowing users to choose the viewpoint from which they observe the scene. The company describes the model as an omni model because it works with all these types of input simultaneously—text, images, video, and 3D data. The idea is that the model stores all inputs in a single shared spatial context. Within this context, each image has its own position in space rather than merely a place in a data queue. Atlas then predicts what comes next while ensuring that it remains consistent with everything it has already seen. It imagines the rest of the scene that is missing from the input.

Pixel-Precise Camera Control

The first group of tasks the model can handle involves generating images with a precisely specified camera. Atlas receives one to six reference images along with a description of the camera movement, treating the camera geometry as a full-fledged input rather than as a textual instruction such as turn left. The output can be a video up to one minute long in 1440p resolution.

World Labs demonstrates this with an example in which it gave the model a single photograph. Atlas used it to create an entire scene, including objects that do not appear in the image at all. It reconstructed the hidden side of a robot and added a strip of grass beside the pool because, based on its knowledge of the world, it assumes that such an element belongs there.

The spatial context also enables a trick that does not work elsewhere. Two unrelated images can be placed into the context at different positions in space, and the model creates a smooth transition between them. It invents a door, corridor, or alcove through which one scene leads into the other.

Reconstructing Real Places from a Few Photos

The second area is the reconstruction of real-world spaces. This is a long-standing computer vision problem: how to infer the view from a new position when only a few images are available. According to the company, Atlas does not require specialized equipment or hundreds of densely captured shots.

What is interesting is how it allows fidelity and imagination to be balanced. The more images the model receives, the less it invents. It can produce faithful reconstructions from just two or three photos, outperforming even models trained exclusively for 3D reconstruction in tests. At the same time, it can handle more than a hundred input images when the goal is to create the most accurate possible copy of a specific location.

The results do not have to be limited to images and videos. Atlas works simultaneously with image frames and depth maps, so it can also generate genuine three-dimensional outputs, such as points in space or so-called Gaussian splats. These can be rendered directly on a device at high resolution and with a smooth frame rate. The same representation is used in Marble, a tool that World Labs offers to the public and whose future versions are expected to be built on Atlas.

Training Robots from Mobile Video

Robotics is one of the areas World Labs is targeting. For robots, simply reconstructing a space is not enough. As a simulated robot moves through the space, Atlas simultaneously calculates the image and depth data that its sensors would perceive at every moment. Both the world and the robot’s view of that world are generated within a single model.

The company recorded two large spaces using mobile video and used 24 frames to reconstruct each of them. Scanning similar locations usually requires expensive and complicated equipment. It then simulated various types of robots moving along different routes within the reconstructed spaces.

The model is most effective at object manipulation. Using a few informal recordings, it helps build a simulation that captures how objects move and interact with one another, including rigid, articulated, and deformable objects. A simulated task can then be varied easily by changing the objects, their positions, the robot’s movement, the lighting, and the background. This creates a large volume of diverse training and testing data.

How Atlas Works

Technically, it is a multimodal autoregressive diffusion transformer. Each of these characteristics serves a specific purpose. Multimodality means that the model processes text, images, camera positions, and depth maps, while treating video as a sequence of frames. The autoregressive component ensures that outputs are generated progressively, each taking the preceding context into account, making the same architecture suitable for many different tasks. The diffusion component gradually removes noise from the output and makes it possible to trade speed for quality depending on how many denoising steps the developer chooses. The transformer provides the entire system with a foundation that efficiently utilizes current hardware.

World Labs emphasizes that it has combined techniques from large language models and video models. Thanks to the autoregressive transformer, Atlas can benefit from techniques used to accelerate language models, such as caching intermediate results or distributing computation across multiple machines. As a latent-space diffusion model, it also takes advantage of improvements developed for image and video generation.

Tests Against Competitors

According to the company, no single test can fully evaluate a model capable of handling so many different tasks. It therefore selected two areas.

For generation with a specified camera, the models received one input photo and one to three cinematic camera movements, such as a pan, dolly shot, or crane shot. Atlas received the camera path in its own native format, while the other models were given verbal descriptions because they do not support any other type of input. People then selected which output better matched the specified movement. Atlas won every comparison, receiving from three-quarters of the votes against MiniMax H3 to more than 90 percent against the FLUX 3 and Seedance 2.5 models. According to World Labs, the gap increases as the camera movement becomes more complex.

The second test concerned 3D reconstruction from a small number of views. The model received a set of images with camera positions and had to determine the corresponding point in space for each pixel. Atlas achieved the lowest average error, outperforming five recent specialized models, including Pi3X, VGGT-Ω 1B, Depth Anything 3, and MapAnything. World Labs states that it recalculated the results of all the compared methods itself to ensure that they all ran under the same conditions.

Source: theinformation.com

Category:AI
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

Kimi K3 Creator Eyes IPO. Moonshot AI Seeks to Raise $3 BillionKimi K3 Creator Eyes IPO. Moonshot AI Seeks to Raise $3 Billion
Moonshot AI, creator of the Kimi K3 model, is planning one of Hong Kong’s largest IPOs in recent years. The company aims to raise $3 billion despite US allegations.
2 min read
4. 9. 2026
New GPT-6 Astra Model: What It Can Do, Where It Fails, and How Much It CostsNew GPT-6 Astra Model: What It Can Do, Where It Fails, and How Much It Costs
GPT-6 Astra promises faster work with computers, office tasks, and code. How convincing are its results, where does it hit its limits, and how much will you pay for it?
6 min read
4. 9. 2026
Solaris doesn’t write code—it renders the screen in real time without HTML, CSS, or JavaScriptSolaris doesn’t write code—it renders the screen in real time without HTML, CSS, or JavaScript
Runway introduces Solaris, which turns the screen into a living environment: it renders the interface frame by frame and responds to prompts and gestures without traditional programming.
6 min read
3. 9. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok