Nvidia's New Cosmos Policy Technology Teaches Robots to Predict the Future and Plan Movements

Nvidia's New Cosmos Policy Technology Teaches Robots to Predict the Future and Plan Movements

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
3. 2. 2026
5 minutes reading
Nvidia's New Cosmos Policy Technology Teaches Robots to Predict the Future and Plan Movements

NVIDIA has introduced Cosmos Policy, a new approach to robotic control built on a world foundation model platform for physical artificial intelligence. Cosmos Policy offers a streamlined way for robots to decide on their actions by adapting large pretrained video models for control and planning tasks.

The key innovation is that Cosmos Policy requires no architectural modifications to the original video model. Instead, it uses a single stage of post-training on robot demonstration data, significantly simplifying the entire process compared to previous methods that required multiple training stages and new architectural components.

How Cosmos Policy works

Traditional robotic systems typically need separate modules for perception, planning, and control. Each of these modules requires large amounts of labeled data and specific tuning for each robot or environment. Cosmos Policy takes a different approach – instead of creating a new control model from scratch, it post-trains a pretrained video model known as Cosmos Predict on robot demonstration data.

The model already understands how the physical world evolves over time because it was trained on extensive video data. During post-training, robot actions, physical states, and task outcomes are processed as part of the model's internal temporal representation. This allows the model to predict not only what the robot should do next, but also what will happen as a result of that action.

Cosmos Policy uses a technique called "latent frame injection". New modalities such as robot proprioception, action sequences, and state values are encoded as new latent frames that are inserted directly into the video model's latent diffusion sequence. This means the model can jointly predict actions, future states, and the expected value of success within a single architecture.

Record-breaking benchmark results

Cosmos Policy achieved impressive results in standard robotics benchmarks. In the LIBERO benchmark, it achieved an average success rate of 98.5% across four task suites, setting a new record. In the Object category specifically, it even achieved a 100% success rate.

LIBERO benchmark
LIBERO benchmark results.

In the RoboCasa benchmark, which includes 24 kitchen manipulation tasks, Cosmos Policy achieved an average success rate of 67.1%, while requiring significantly fewer demonstrations for training than competing methods – only 50 demonstrations compared to 300 or more for other approaches.

RoboCasa benchmark
RoboCasa benchmark results.

In real-world experiments with the ALOHA bimanual robot, Cosmos Policy outperformed all compared methods, including advanced vision-language-action models such as π₀.₅ and OpenVLA-OFT+. It achieved the highest average score of 93.6% on challenging tasks requiring long-horizon, high-precision manipulation.

Real-world results of the ALOHA robot.
Real-world results of the ALOHA robot.

Planning using a world model

One of the key features of Cosmos Policy is its ability to perform planning at inference time. Instead of producing only the immediate next action, the model can generate and evaluate multiple candidate action sequences. By predicting the future outcomes and expected rewards of these sequences, the robot can select actions that are more likely to succeed over a longer time horizon.

Cosmos Policy uses best-of-N sampling – it samples multiple action proposals from the model, uses a planning model to predict the future state and value for each proposal, and selects and executes the action that leads to the predicted state with the highest value. For greater accuracy, the model uses ensemble predictions – it queries the world model three times per action and the value function five times per future state, resulting in a total of fifteen value predictions for each action proposal.

In challenging real-world manipulation tasks, model-based planning led to an average score increase of 12.5 percentage points across the two most challenging tasks. Qualitatively, the researchers found that the fine-tuned planning model predicts future states more accurately and can plan more effectively, thereby avoiding mistakes made by the baseline Cosmos Policy.

Part of the NVIDIA Cosmos ecosystem

Cosmos Policy is part of the broader NVIDIA Cosmos platform, which focuses on building universal world models for robots and autonomous systems. The platform includes several key components:

Cosmos Predict generates up to 30 seconds of high-quality video from multimodal prompts and predicts future states of dynamic environments for planning by robots and AI agents.

Cosmos Transfer accelerates synthetic data generation across different environments and lighting conditions, transforming 3D or spatial inputs from simulation frameworks such as CARLA or NVIDIA Isaac Sim into fully controlled, high-quality video.

Cosmos Reason is a multimodal vision-language model that enables robots and vision AI agents to reason like humans, using prior knowledge, an understanding of physics, and common sense to understand the real world.

Technical details and availability

Cosmos Policy was developed by a team of researchers from NVIDIA and Stanford University, including Moo Jin Kim, Yihuai Gao, Tsung-Yi Lin, Yen-Chen Lin, and others. The model is based on Cosmos-Predict2-2B, a latent video diffusion model with 2 billion parameters.

All Cosmos models, including Cosmos Policy, are available under the NVIDIA Open Model License and are available on GitHub and Hugging Face. NVIDIA also provides the Cosmos Cookbook – a practical guide featuring step-by-step workflows, technical recipes, and concrete examples for building, customizing, and deploying world foundation models.

For developers who want to get started with Cosmos Policy, NVIDIA offers several options: direct access to the models and code on GitHub, trying the models in the hosted catalog, or using the practical recipes in the Cosmos Cookbook.

Cosmos Policy represents a significant advancement in robot learning by combining the power of pretrained video models with efficient post-training on robotic data, achieving state-of-the-art results while maintaining simplicity and flexibility.

Source: interestingengineering.com

Category:Robotics
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

China Unveils First Robotic Centaur. It Will Work Where Humans DieChina Unveils First Robotic Centaur. It Will Work Where Humans Die
At a Shanghai exhibition center this July, a machine appeared that at first glance defied classification. Shanghai-based Run Robotics unveiled a robot at the WAIC 2026 conference that it calls
3 min read
27. 7. 2026
Not for homes, but robots for industry. English firm Humanoid reaches billion-dollar valuationNot for homes, but robots for industry. English firm Humanoid reaches billion-dollar valuation
London-based robotics company Humanoid, officially known as SKL Robotics, has joined the ranks of so-called unicorns—young companies valued at more than one billion dollars. According to...
6 min read
20. 7. 2026
A Breakthrough in Robotics: Robots Get Artificial Hands That Work Like Ours!A Breakthrough in Robotics: Robots Get Artificial Hands That Work Like Ours!
California-based 1X from Palo Alto has unveiled new hands for its NEO robot. They have 25 degrees of freedom, are tendon-driven like human hands and, according to the manufacturer, approach human dexterity, strength and reliability.
5 min read
14. 7. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok