MiniMax H3 · Hailuo 03 multimodal video

MiniMax H3 AI Video Generator

Create video from text, animate a first frame with an optional final frame, or combine image, video and audio references. Generate 4–15 second clips at 768p or 2K with native audio.

Input

Generation mode
0 / 5000

Explain the shot order and the role of every reference, including what must remain consistent.

Aspect ratio
Video duration
Resolution

Output

Your video will appear here

Choose a MiniMax H3 mode, add a prompt and any required references, then submit the task.

The model, explained

What is MiniMax H3?

MiniMax H3, also known as Hailuo 03, is a multimodal video model that understands text, images, video and audio in one creative context.

It supports text-to-video, first-and-last-frame image animation, and reference-to-video workflows with 4–15 second output at 768p or 2K.

FramePack submits each request asynchronously, tracks its progress and privately stores the completed video for preview and download.

Unified audiovisual creation

Three MiniMax H3 workflows

Move from a prompt or mixed references to a complete video with synchronized sound.

Detailed scene direction

Direct characters, actions, shot order, camera movement, dialogue, music and sound effects in one prompt.

Multimodal references

Use opening and closing frames or combine multiple image, video and audio references.

768p and 2K output

Choose 4–15 seconds, supported framing and the resolution that fits the final use.

Tracked private delivery

Follow the task while it renders, then preview or download the privately stored result.

Three focused steps

How to create a MiniMax H3 video

Choose a workflow, direct every input clearly and review the finished clip.

  1. 01

    Choose the starting point

    Start from text, a first and optional last frame, or a set of multimodal references.

  2. 02

    Direct and configure

    Write the prompt, assign each reference a role, then set duration, resolution and framing.

  3. 03

    Generate and download

    Submit the asynchronous task, follow its status, and download the stored video when ready.

MiniMax H3 prompting tips

Structured directions help visual and audio elements stay coherent across the clip.

  • Describe events in time order and include camera and sound cues beside the related action.
  • State whether each file controls identity, motion, style, voice, music or another specific element.
  • Choose the final aspect ratio before composing the scene and reserve 2K for higher-detail delivery.
  • List details that must stay unchanged separately from the edits or transformations you want.

Questions, answered

MiniMax H3 FAQ

What can MiniMax H3 generate?

It creates 4–15 second audiovisual clips from text, frames, or combined image, video and audio references at 768p or 2K.

Which reference files are supported?

Image-to-video accepts one or two images. Reference-to-video accepts up to nine images, three MP4 or MOV videos and three MP3 or WAV audio files.

Does MiniMax H3 generate sound?

Yes. The model can create native stereo audio, and reference-to-video can use uploaded audio to guide voice, music or sound.

How is the result delivered?

After the generation succeeds, FramePack privately stores the video and shows authenticated preview and download links.

Create a complete audiovisual shot

Combine a clear prompt with focused references and generate with MiniMax H3.

Create with MiniMax H3