Odyssey has just launched Odyssey-2, a new interactive video model that creates streamed AI video at 20 frames per second. Users can shape and control videos lasting several minutes using text commands while exploring the scene. This system differs from other video models that take minutes to produce short clips. Odyssey-2 starts streaming video immediately, with new frames appearing every 50 milliseconds.
The model generates video segments without prior planning. Each new frame is based on what has already happened and what users enter in real time. This means the video evolves dynamically, without a fixed ending. For example, as the model learns physics and dynamics from video data, it can simulate realistic behavior, such as waves moving across water or changes in light on surfaces.
Users control how the video develops through natural-language commands in a chat window while the video is running. The AI continuously adapts to each input, creating a sense of live interaction. Odyssey-2 is based on a causal and autoregressive approach, where each frame is created solely from previous frames and user actions, without knowledge of the future.

Odyssey-2 Speed and Technical Details
Odyssey-2 achieves speeds of up to 30 frames per second, with each frame generated in less than 50 milliseconds. This speed transforms the entire experience because, instead of waiting for a short clip, you get an instant stream that responds to your commands. The model has been optimized at the architecture, data pipeline, and inference stack levels to balance speed, quality, and responsiveness.
The system learns complex physical phenomena from decades of video data. For example, when generating ocean waves, it estimates the surface slope, curvature, and velocity field from previous frames, then predicts the next movement—the crest advances, troughs fill, foam drifts, and the wave bends around a rock. All of this happens in real time, making the model an implicit world simulator.
Odyssey-2 streams videos longer than five minutes while maintaining environmental coherence. It uses clusters of Nvidia H100 GPUs for performance, with operating costs of $1–2 per user hour. Data is collected using a 360-degree camera in a backpack to produce realistic outputs.
Comparison with Competitors
Odyssey-2 feels different from other AI video platforms such as Veo or Sora. It is more of a hybrid world generator than simply a clip creator. Although the quality may not be as impressive as that of other models, real-time operation and open-ended exploration offer enormous potential for new content experiences.
The model focuses on open-ended interactivity, where an action at any moment changes every possible future. This enables continuous video streaming that listens, adapts, and responds. Odyssey-2 builds on Odyssey-1, which focused on navigation and long-term memory, and expands it with text commands and, soon, audio commands as well.
Applications and the Future of Odyssey-2
Odyssey-2 opens the door to new applications in gaming, film, education, social media, advertising, training, and simulations. Imagine walking through an old photograph and exploring memories, or taking a guided tour of an ancient civilization. The model integrates with tools such as Unreal Engine, Blender, and Adobe After Effects for creators.
This system moves from fixed media to emergent media, where video responds to the user. It is still at an early stage, but Odyssey-2 delivers an experience similar to conversing with a language model—type something, and the video responds immediately. The quality is still rough and sometimes unstable, but development continues toward greater realism and interactivity.



