Natural-language direction
Describe the scene, action, camera, style, lighting and requested transformation in one prompt.
Gemini Omni multimodal video
Create a video from a prompt and visual references, or transform an existing clip with natural-language direction. Choose the format and follow the asynchronous result online.
Output
Your video will appear here
Write a prompt, add optional reference media and submit a Gemini Omni task.
The model, explained
Gemini Omni is a multimodal AI video generator that can interpret a written prompt together with reference images and an optional source video.
Use the same workspace to create a new clip from text and images or transform existing footage while describing the elements that should change and remain consistent.
The page sends the request to Kie as an asynchronous task, follows its progress, and privately stores the completed video for authenticated preview and download.
Multimodal video direction
Combine natural-language direction with visual context, then choose the output settings that fit your delivery channel.
Describe the scene, action, camera, style, lighting and requested transformation in one prompt.
Add images for visual guidance and optionally one source video for transformation or editing.
Choose horizontal or vertical framing, 720p, 1080p or 4K output, and a supported duration when no source video is present.
Monitor the Kie task and access the finished video through private preview and download links.
Three focused steps
Give the model a clear instruction, add only the visual context it needs, and review the completed result.
Define the subject, action, visual treatment and camera behavior, including what must stay consistent during an edit.
Upload reference images or an optional source video, then choose aspect ratio, resolution and duration when available.
Submit the multimodal task, follow its status, and preview or download the privately stored output.
Use visual references selectively and make every requested transformation explicit.
Questions, answered
It can create a video from a prompt and reference images, or transform one uploaded video with natural-language direction.
The page supports up to seven media units. Each image uses one unit and an optional source video uses two, leaving room for up to five images beside that video.
Without a source video, you can choose 4, 6, 8 or 10 seconds. When a source video is included, Gemini Omni determines the output duration automatically.
FramePack stores the completed result privately and exposes authenticated preview and download links after the Kie task succeeds.
Combine a prompt with focused visual references and create your next clip with Gemini Omni.
Create with Gemini Omni