Meta and Its V-JEPA 2: AI That Understands the World Through Video

Meta and Its V-JEPA 2: AI That Understands the World Through Video

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
12. 6. 2025
4 minutes reading · 8 views
Meta and Its V-JEPA 2: AI That Understands the World Through Video

Meta and Its V-JEPA 2: AI That Understands the World Through Video

Meta AI introduced V-JEPA 2, an advancement in world models for AI agents and robotics that represents a significant step forward in understanding, predicting, and planning actions in complex and unfamiliar environments. This model, with 1.2 billion parameters, is trained primarily on raw video data and represents a fundamental shift in embodied AI. V-JEPA 2 builds on the foundations of the Joint Embedding Predictive Architecture (JEPA) and introduces capabilities that approach the human way of understanding the physical world through visual information.

The model uses an innovative two-stage training process that begins with self-supervised pretraining on more than one million hours of video footage and one million images without human labels. This approach allows the model to capture patterns of physical interactions in the world and create a robust representation of reality. Subsequent action-conditioned fine-tuning uses a relatively small dataset of approximately 62 hours of robotic control data, enabling the model to factor in agent actions when predicting outcomes. This combination of broad-based learning and specific fine-tuning creates a system capable of effective planning and closed-loop control, even with limited domain-specific examples.

Zero-Shot Planning and Control

One of the most remarkable features of V-JEPA 2 is its ability to perform zero-shot planning and robotic control, meaning that the model can generalize to new tasks and environments without extensive retraining. This feature significantly increases the flexibility and adaptability of robotic systems, marking an important milestone in autonomous systems. The model can work with vision-based representations of goals and sequences of visual subgoals for more complex tasks, allowing robots to perform sophisticated operations such as “pick and place” with minimal prior configuration.

Its zero-shot generalization capability stems from the broad representation of the physical world that V-JEPA 2 develops during pretraining. When the model encounters a new task or environment, it can use its existing knowledge of physical laws, object interactions, and causal relationships to devise effective strategies. This approach represents a dramatic shift from traditional methods, which require extensive training for each new domain or task. V-JEPA 2 therefore paves the way for more versatile robotic systems that can quickly adapt to the changing demands of the real world.

Applications in Robotics and Assistive Technologies

In robotics, V-JEPA 2 enables robots to perform complex tasks using vision-based representations of goals and sequences of visual subgoals for more sophisticated operations. The model can generate detailed plans for manipulating objects, navigating through space, and coordinating multiple subtasks within larger operations. This capability is especially valuable in industrial applications, where robots must work with varying objects and changing conditions in the production environment.

Beyond traditional robotics, V-JEPA 2 has significant potential in assistive technologies, where it could enhance wearable assistants for real-time environmental recognition and navigational support. These applications could be particularly beneficial for people with visual impairments or other disabilities, providing them with sophisticated tools for navigating and interacting with the world around them. The model can analyze complex visual scenes and provide users with relevant information about their surroundings, identify obstacles, recognize objects, and navigate safely through the environment.

New Benchmarks for Evaluating World Models

Alongside V-JEPA 2, Meta introduced three new benchmarks designed to accelerate research into physical reasoning and world modeling. These benchmarks are specifically designed to evaluate the ability of AI models to learn and reason about the world using video data. The goal is to provide the research community with standardized tools for comparing and improving the physical reasoning capabilities of world models. The benchmarks cover various aspects of world modeling, including predicting physical interactions, understanding causal relationships, and the ability to generalize to new situations.

These benchmarks are an important step toward systematically evaluating progress in embodied AI and provide objective metrics for comparing different approaches. Standardized evaluation will allow researchers to identify the strengths and weaknesses of different models and direct research efforts toward areas with the greatest potential for improvement. The benchmarks also facilitate research reproducibility and enable fair comparisons between different laboratories and organizations.

Open Science and a Community-Based Approach

Meta has adopted an open-science approach by releasing V-JEPA 2 and its benchmarks to the broader community in order to foster innovation and progress in embodied AI and world modeling. By enabling the research community to build on this work, Meta aims to drive progress toward advanced machine intelligence that learns and adapts as efficiently as humans do. This community-based approach reflects a growing trend in AI research, where the open sharing of resources and tools accelerates collective progress in the field.

Releasing V-JEPA 2 as open source will allow researchers around the world to experiment with the model, adapt it for specific applications, and contribute to its continued development. This approach could lead to faster discovery of new applications and improved model performance through distributed research efforts. Meta is thus demonstrating a commitment to advancing the entire field of AI rather than merely developing proprietary technology, which could have far-reaching positive effects on the pace of innovation in robotics and autonomous systems.

Advertisement

Content created with help from UpTier.

SEO and GEO on autopilot. UpTier’s multi-agent systems write and optimize content for search engines and AI answers.

Discover UpTier ↗

Category:AI
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

OpenAI gives Codex reusable cloud workspaces accessible from any deviceOpenAI gives Codex reusable cloud workspaces accessible from any device
Codex gains reusable cloud development environments, alongside voice controls in its CLI, code reviews in the ChatGPT desktop app and cloud-based security tools.
2 min read
2. 10. 2026
Amazon releases Strands Decider 2B for AI workflow decisionsAmazon releases Strands Decider 2B for AI workflow decisions
Strands Decider 2B selects from predefined options and returns a confidence score. The fully open-source model is available now and small enough to run locally.
2 min read
1. 10. 2026
OpenAI says it disrupted a campaign to extract hidden model reasoningOpenAI says it disrupted a campaign to extract hidden model reasoning
OpenAI reported a coordinated effort to extract protected model reasoning and said it closed an extraction pathway. It attributed the main cluster of activity to individuals associated with Moonshot AI, the developer of Kimi.
3 min read
1. 10. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok