Detailed scene direction
Direct characters, actions, shot order, camera movement, dialogue, music and sound effects in one prompt.
MiniMax H3 · Hailuo 03 multimodal video
Create video from text, animate a first frame with an optional final frame, or combine image, video and audio references. Generate 4–15 second clips at 768p or 2K with native audio.
Output
Your video will appear here
Choose a MiniMax H3 mode, add a prompt and any required references, then submit the task.
The model, explained
MiniMax H3, also known as Hailuo 03, is a multimodal video model that understands text, images, video and audio in one creative context.
It supports text-to-video, first-and-last-frame image animation, and reference-to-video workflows with 4–15 second output at 768p or 2K.
FramePack submits each request asynchronously, tracks its progress and privately stores the completed video for preview and download.
Unified audiovisual creation
Move from a prompt or mixed references to a complete video with synchronized sound.
Direct characters, actions, shot order, camera movement, dialogue, music and sound effects in one prompt.
Use opening and closing frames or combine multiple image, video and audio references.
Choose 4–15 seconds, supported framing and the resolution that fits the final use.
Follow the task while it renders, then preview or download the privately stored result.
Three focused steps
Choose a workflow, direct every input clearly and review the finished clip.
Start from text, a first and optional last frame, or a set of multimodal references.
Write the prompt, assign each reference a role, then set duration, resolution and framing.
Submit the asynchronous task, follow its status, and download the stored video when ready.
Structured directions help visual and audio elements stay coherent across the clip.
Questions, answered
It creates 4–15 second audiovisual clips from text, frames, or combined image, video and audio references at 768p or 2K.
Image-to-video accepts one or two images. Reference-to-video accepts up to nine images, three MP4 or MOV videos and three MP3 or WAV audio files.
Yes. The model can create native stereo audio, and reference-to-video can use uploaded audio to guide voice, music or sound.
After the generation succeeds, FramePack privately stores the video and shows authenticated preview and download links.
Combine a clear prompt with focused references and generate with MiniMax H3.
Create with MiniMax H3