xAI has released the latest version of its video generation model. Grok Imagine Video 1.5 takes your image, you describe what should happen in it, and the model turns it into a short video with motion, sound, and spoken dialogue. Most importantly, it does so quickly. While its predecessor previously needed more than forty seconds, it can now do it in roughly twenty-five.
Speed, sound, and better physics
Let's start with what people will notice first: speed. The Grok Imagine Video 1.5 Fast variant has nearly doubled generation speed. It can produce a six-second video in 720p resolution in around twenty-five seconds. The previous model needed more than forty seconds. Why does this matter? When you're creating content and waiting a minute for each clip, the work drags on. This significantly speeds up your workflow.
The model generates not only visuals but also sound. Sound effects, ambient noise, and dialogue are created in the same pass and precisely match what is happening on screen. Speech is clearer and better synchronized with lip movements. You describe the scene, and the model provides both the visuals and the corresponding audio track. This eliminates the need to edit audio in a separate program.
Older video generation models often had problems with characters and objects becoming distorted during a clip. Grok Imagine Video 1.5 keeps motion consistent throughout the entire shot. There are fewer strange distortions, and objects have believable weight and inertia. When something moves, it looks more natural.
Along with the model, xAI has added several features that make the workflow easier. You can organize your creations into projects, which you will find in the left panel. You can also run multiple tasks at once, so you do not have to wait for one video to finish before starting another. Simply enter several prompts and let them run simultaneously. And when you are looking for one specific clip, you can search the library instead of scrolling endlessly.
How to use it and what you need
You provide the model with an input image, describe the motion, and select the resolution and duration. xAI provides sample code where you simply enter a prompt, a link to the image, the video duration in seconds, and the resolution. Then you just wait for the result.
The model is available through the xAI interface as grok-imagine-video-1.5 You can also try it directly at grok.com/imagine or in the iOS and Android apps, which use the aforementioned fast variant.
Where to find the model and how much it costs
In addition to direct access through xAI, Grok Imagine is also available on the Kie.ai platform, which brings together multiple video generation models in one place. There, you can choose from four input types, ranging from image-to-video and text-to-video to text-to-image.
Kie.ai calculates the price based on the duration of the resulting video. Output in 480p costs 1.6 credits per second, while 720p resolution costs 3 credits per second. Various generation modes are available, including normal, fun, and the more expressive “Spicy” mode. However, Spicy mode does not work for videos created from your own uploaded image and automatically switches to normal.
You can also choose from aspect ratios ranging from vertical 9:16 to widescreen 16:9 and set the clip duration as needed.



