Meta and Its V-JEPA 2: AI That Understands the World Through Video

Meta and Its V-JEPA 2: AI That Understands the World Through Video

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
12. 6. 2025
4 minutes reading
Meta and Its V-JEPA 2: AI That Understands the World Through Video

Meta and Its V-JEPA 2: AI That Understands the World Through Video

Meta AI introduced V-JEPA 2, an advancement in world models for AI agents and robotics that represents a significant step forward in understanding, predicting, and planning actions in complex and unfamiliar environments. This model, with 1.2 billion parameters, is trained primarily on raw video data and represents a fundamental shift in embodied AI. V-JEPA 2 builds on the foundations of the Joint Embedding Predictive Architecture (JEPA) and introduces capabilities that approach the human way of understanding the physical world through visual information.

The model uses an innovative two-stage training process that begins with self-supervised pretraining on more than one million hours of video footage and one million images without human labels. This approach allows the model to capture patterns of physical interactions in the world and create a robust representation of reality. Subsequent action-conditioned fine-tuning uses a relatively small dataset of approximately 62 hours of robotic control data, enabling the model to factor in agent actions when predicting outcomes. This combination of broad-based learning and specific fine-tuning creates a system capable of effective planning and closed-loop control, even with limited domain-specific examples.

Zero-Shot Planning and Control

One of the most remarkable features of V-JEPA 2 is its ability to perform zero-shot planning and robotic control, meaning that the model can generalize to new tasks and environments without extensive retraining. This feature significantly increases the flexibility and adaptability of robotic systems, marking an important milestone in autonomous systems. The model can work with vision-based representations of goals and sequences of visual subgoals for more complex tasks, allowing robots to perform sophisticated operations such as “pick and place” with minimal prior configuration.

Its zero-shot generalization capability stems from the broad representation of the physical world that V-JEPA 2 develops during pretraining. When the model encounters a new task or environment, it can use its existing knowledge of physical laws, object interactions, and causal relationships to devise effective strategies. This approach represents a dramatic shift from traditional methods, which require extensive training for each new domain or task. V-JEPA 2 therefore paves the way for more versatile robotic systems that can quickly adapt to the changing demands of the real world.

Applications in Robotics and Assistive Technologies

In robotics, V-JEPA 2 enables robots to perform complex tasks using vision-based representations of goals and sequences of visual subgoals for more sophisticated operations. The model can generate detailed plans for manipulating objects, navigating through space, and coordinating multiple subtasks within larger operations. This capability is especially valuable in industrial applications, where robots must work with varying objects and changing conditions in the production environment.

Beyond traditional robotics, V-JEPA 2 has significant potential in assistive technologies, where it could enhance wearable assistants for real-time environmental recognition and navigational support. These applications could be particularly beneficial for people with visual impairments or other disabilities, providing them with sophisticated tools for navigating and interacting with the world around them. The model can analyze complex visual scenes and provide users with relevant information about their surroundings, identify obstacles, recognize objects, and navigate safely through the environment.

New Benchmarks for Evaluating World Models

Alongside V-JEPA 2, Meta introduced three new benchmarks designed to accelerate research into physical reasoning and world modeling. These benchmarks are specifically designed to evaluate the ability of AI models to learn and reason about the world using video data. The goal is to provide the research community with standardized tools for comparing and improving the physical reasoning capabilities of world models. The benchmarks cover various aspects of world modeling, including predicting physical interactions, understanding causal relationships, and the ability to generalize to new situations.

These benchmarks are an important step toward systematically evaluating progress in embodied AI and provide objective metrics for comparing different approaches. Standardized evaluation will allow researchers to identify the strengths and weaknesses of different models and direct research efforts toward areas with the greatest potential for improvement. The benchmarks also facilitate research reproducibility and enable fair comparisons between different laboratories and organizations.

Open Science and a Community-Based Approach

Meta has adopted an open-science approach by releasing V-JEPA 2 and its benchmarks to the broader community in order to foster innovation and progress in embodied AI and world modeling. By enabling the research community to build on this work, Meta aims to drive progress toward advanced machine intelligence that learns and adapts as efficiently as humans do. This community-based approach reflects a growing trend in AI research, where the open sharing of resources and tools accelerates collective progress in the field.

Releasing V-JEPA 2 as open source will allow researchers around the world to experiment with the model, adapt it for specific applications, and contribute to its continued development. This approach could lead to faster discovery of new applications and improved model performance through distributed research efforts. Meta is thus demonstrating a commitment to advancing the entire field of AI rather than merely developing proprietary technology, which could have far-reaching positive effects on the pace of innovation in robotics and autonomous systems.

Category:AI
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

Altman Announced the Singularity Days After His Models Escaped the Lab on Their OwnAltman Announced the Singularity Days After His Models Escaped the Lab on Their Own
OpenAI chief Sam Altman declared on the Relentless podcast that humanity has already entered the singularity. “We’re like, in the singularity now,” he said verbatim. For decades, the term belonged more to science-fiction literature
6 min read
28. 7. 2026
AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.
Since April, Australian radio has been playing a dance remake of Madonna’s hit Like a Prayer on repeat. Released by Queensland DJ Josh Fawaz, it tops the radio airplay chart and has 35 million Spotify streams.
6 min read
28. 7. 2026
Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?
A first-person shooter that runs directly in the browser, with its own physics and eleven separate code modules. Around 55,000 lines in total, split across eleven subsystems and built on Thr
4 min read
28. 7. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok