Google Has a New, Better AI Video Generator. All You Need Is Text, an Image, Video, or Audio

Google Has a New, Better AI Video Generator. All You Need Is Text, an Image, Video, or Audio

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
21. 5. 2026
3 minutes reading
Google Has a New, Better AI Video Generator. All You Need Is Text, an Image, Video, or Audio

    Google DeepMind has introduced a new model called Gemini Omni, which can create videos from virtually any input material. Text, a photo, audio, or existing video footage. Any of these can serve as the starting point. The result is always a video. The model is initially being released as Gemini Omni Flash and is available in the Gemini app, Google Flow, and YouTube Shorts. Google is deploying it as the direct successor to the Veo model, which previously handled video generation in the Gemini app.

    Edit easily via chat using different inputs

    What sets Gemini Omni apart from previous tools? The way you edit. You do not work with tracks and layers as you would in a traditional editor. You simply write what you want to change. Want to move a violinist to a different setting? Write it. Then want to hide the violin? Write it. And then change the camera angle to an over-the-shoulder shot. Each edit builds on the previous one, the scene remains consistent, and the characters retain their appearance. The system remembers the context of the entire sequence.

    Gemini Omni can handle multi-step edits while preserving the image’s physical logic. Liquids behave like liquids. A marble rolls the way it should. Google describes these capabilities as an intuitive understanding of forces such as gravity, kinetic energy, and fluid dynamics.

    One of the most interesting things Google demonstrated when introducing the model was the combination of different input types in a single output. A user can attach a video capturing movement, a photo of a character, and a music track. Gemini Omni combines them into a single video in which the character from the photo moves in time with the music and follows the style of the reference footage. The inputs are not combined mechanically; the model looks for narrative logic.

    For now, direct audio input works only with voice recordings. Google plans to make other types of audio inputs available gradually.

    Another interesting feature is sketch input. Sketch a fish, a bird, or a dandelion on paper, take a photo of it, and Gemini Omni will turn it into a realistic video. The movement indicated in the drawing serves as a guide for the movement in the resulting footage. The drawing itself does not appear in the video. Replacing characters or objects works in a similar way. You attach a photo of a character and tell the model, "turn me into this character." The resulting character adopts the movement, expression, and dialogue from the original footage.

    Google emphasizes that the model draws on Gemini’s knowledge base, which includes history, science, mathematics, and cultural context. In the demonstrations, this means, for example, a video explaining protein folding or an alphabet series featuring unusual objects for each letter, all automatically synchronized with music and captions. So the model not only generates images but also understands what it is depicting.

    Gemini Omni Flash is available to users aged 18 and over with a Google AI Plus, Pro, or Ultra subscription. The service works in all languages and markets where the Gemini app is available. Some features, such as video or avatar editing, may be restricted in certain countries.

    Videos created with Gemini are marked with an invisible SynthID watermark and contain metadata based on the C2PA standard, which makes it possible to verify the origin of the content. Verification will soon be available directly in the Chrome browser and Google Search.

    Category:AI
    Did you enjoy this article?
    Discover more interesting posts on our blog
    Back to blog

    Related posts

    Altman Announced the Singularity Days After His Models Escaped the Lab on Their OwnAltman Announced the Singularity Days After His Models Escaped the Lab on Their Own
    OpenAI chief Sam Altman declared on the Relentless podcast that humanity has already entered the singularity. “We’re like, in the singularity now,” he said verbatim. For decades, the term belonged more to science-fiction literature
    6 min read
    28. 7. 2026
    AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.
    Since April, Australian radio has been playing a dance remake of Madonna’s hit Like a Prayer on repeat. Released by Queensland DJ Josh Fawaz, it tops the radio airplay chart and has 35 million Spotify streams.
    6 min read
    28. 7. 2026
    Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?
    A first-person shooter that runs directly in the browser, with its own physics and eleven separate code modules. Around 55,000 lines in total, split across eleven subsystems and built on Thr
    4 min read
    28. 7. 2026
    Přihlaste se k odběru našeho newsletteru
    Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
    CodedTrip

    Operated by CodedTrip LLC, USA.

    YouTube
    TikTok