Hailuo 3
MiniMax · AI videoHailuo 3 (MiniMax H3) is MiniMax's video generation model. It turns a text prompt and optional first/last frames into short clips, with native audio generation.
It also accepts an audio URL for lip-synced talking-head clips: pair a portrait (or scene start frame) with speech audio so mouth motion follows the track.
The API accepts 4–15 second durations and common aspect ratios including 9:16 and 16:9. First-frame and last-frame images can condition the start and end of the clip when you are not driving from reference audio.
On VidMachine, scene generation uses Protoface at 480p (3 credits/s) or 768p (8 credits/s). Standard AI Video uses image-to-video; AI Influencer and lip-sync playground send Protoface `audio` plus a `reference` image.
Key features and benefits
Text-to-video and image-to-video
Describe the clip with a prompt, or condition generation on a first frame (and optionally a last frame) so motion starts from your key art.
Reference audio lip sync
Provide speech audio with a portrait or scene image. VidMachine sends Protoface fields audio (URL) and reference (image). First/last-frame mode is omitted when audio is present.
Selectable resolution
Choose 480p or 768p in project settings. Higher resolution costs more credits per second of output.
Native audio
Optional audio generation produces sound alongside the video in one pass. VidMachine always enables it; with a reference audio URL the speech track guides lip sync.
Duration
Clips run 4–15 seconds per generation. Reference audio clips should also stay within about 4–15 seconds.
Technical specifications
ProviderMiniMax (Hailuo 3 / H3)
Primary inputsText prompt; optional first-frame and last-frame images; or reference image + audio for lip sync
Resolution on VidMachine480p or 768p (project model setting)
Aspect ratio21:9, 16:9, 4:3, 1:1, 3:4, 9:16
Duration4–15 seconds
AudioNative audio generation; optional Protoface audio URL for lip sync
Pricing on VidMachine3 credits/s at 480p; 8 credits/s at 768p
Use cases and applications
Use Hailuo 3 for cinematic shorts, product motion from stills, and 9:16 social clips when you want MiniMax's Hailuo motion profile.
Use it for AI Influencer talking-head segments and the lip-sync playground when you have a portrait plus speech audio.
First-and-last-frame control is useful when you already have start and end keyframes and want the model to interpolate motion between them (without reference audio).
Why this model
Pick Hailuo 3 when you want MiniMax's latest Hailuo video stack with native audio, optional lip sync via audio URL, and a clear 480p vs 768p price tradeoff.
Compare credit cost and look with other models in your project settings before committing a long export.
How VidMachine uses it
For AI Video scenes, VidMachine generates clips via image-to-video through Protoface (minimax/minimax-h3), using each scene's starting frame and motion prompt. Duration is clamped to the model's 4–15 second range. Optional last-frame images are sent when available.
For AI Influencer segments and lip-sync playground runs, VidMachine sends the portrait or scene start frame as reference plus the speech clip as audio (public URL). First/last-frame fields are omitted in that mode.
Choose 480p or 768p in project settings. Billing is 3 credits per second at 480p and 8 credits per second at 768p.
What you should know
What resolutions does VidMachine use?
480p at 3 credits per second and 768p at 8 credits per second. 2K is not selectable.
Does Hailuo 3 support last-frame control?
Yes for standard AI Video scenes. Protoface accepts first_frame and last_frame. VidMachine sends the scene start frame and, when present, the scene last frame. Last-frame is not used together with lip-sync audio.
Is Hailuo 3 available for AI Influencer lip-sync?
Yes. Select MiniMax H3 (480p or 768p) on influencer_video_model_priority or in the lip-sync playground. Each scene uses a start frame (or portrait) as reference plus scene speech as Protoface audio.