Google DeepMind has introduced a new model called Gemini Omni, which can create videos from virtually any input material. Text, a photo, audio, or existing video footage. Any of these can serve as the starting point. The result is always a video. The model is initially being released as Gemini Omni Flash and is available in the Gemini app, Google Flow, and YouTube Shorts. Google is deploying it as the direct successor to the Veo model, which previously handled video generation in the Gemini app.
We’re dropping Gemini Omni: our first step towards a model that can create anything from anything - starting with video.
— Google DeepMind (@GoogleDeepMind) May 19, 2026
It combines Gemini’s intelligence with our generative media systems - representing a leap forward in world understanding, multimodality, and editing 🧵 pic.twitter.com/GAtqzr0VIV
Edit easily via chat using different inputs
What sets Gemini Omni apart from previous tools? The way you edit. You do not work with tracks and layers as you would in a traditional editor. You simply write what you want to change. Want to move a violinist to a different setting? Write it. Then want to hide the violin? Write it. And then change the camera angle to an over-the-shoulder shot. Each edit builds on the previous one, the scene remains consistent, and the characters retain their appearance. The system remembers the context of the entire sequence.
Gemini Omni can handle multi-step edits while preserving the image’s physical logic. Liquids behave like liquids. A marble rolls the way it should. Google describes these capabilities as an intuitive understanding of forces such as gravity, kinetic energy, and fluid dynamics.
One of the most interesting things Google demonstrated when introducing the model was the combination of different input types in a single output. A user can attach a video capturing movement, a photo of a character, and a music track. Gemini Omni combines them into a single video in which the character from the photo moves in time with the music and follows the style of the reference footage. The inputs are not combined mechanically; the model looks for narrative logic.
For now, direct audio input works only with voice recordings. Google plans to make other types of audio inputs available gradually.
Another interesting feature is sketch input. Sketch a fish, a bird, or a dandelion on paper, take a photo of it, and Gemini Omni will turn it into a realistic video. The movement indicated in the drawing serves as a guide for the movement in the resulting footage. The drawing itself does not appear in the video. Replacing characters or objects works in a similar way. You attach a photo of a character and tell the model, "turn me into this character." The resulting character adopts the movement, expression, and dialogue from the original footage.
Google emphasizes that the model draws on Gemini’s knowledge base, which includes history, science, mathematics, and cultural context. In the demonstrations, this means, for example, a video explaining protein folding or an alphabet series featuring unusual objects for each letter, all automatically synchronized with music and captions. So the model not only generates images but also understands what it is depicting.
Gemini Omni Flash is available to users aged 18 and over with a Google AI Plus, Pro, or Ultra subscription. The service works in all languages and markets where the Gemini app is available. Some features, such as video or avatar editing, may be restricted in certain countries.
Videos created with Gemini are marked with an invisible SynthID watermark and contain metadata based on the C2PA standard, which makes it possible to verify the origin of the content. Verification will soon be available directly in the Chrome browser and Google Search.



