Generate picture and sound in one pass with MiniMax H3.
MiniMax H3 is the model Hailuo AI leads with, and the one reviewers usually call Hailuo 3.0. It reads text, images, video and audio as a single context, generates picture and 32 kHz stereo sound in the same pass, and takes edit instructions against footage you already have.
Capabilities and specifications on this page come from MiniMax’s H3 announcement, its open-platform pricing page and the Hailuo AI product site. Prices and licence terms are set by the vendor and change without notice.
What MiniMax H3 renders
These frames are MiniMax’s own H3 outputs, published with the model. The set is unusually typographic — most of them carry large lettering — and lettering is the part of a render that normally falls apart.
Example images belong to MiniMax and are reproduced with credit: MiniMax H3 announcement
How MiniMax describes H3
MiniMax positions H3 as an omni-modal system rather than a text-to-video model: it reads text, images, video and audio as one context and produces a finished audiovisual scene. Four claims carry that pitch.
One context, four modalities
A single request can carry up to 12 files — nine images, three video clips and three audio clips — and each reference gets a declared role: this image locks the face, this clip is the motion reference, this track sets the pacing.
Sound is generated, not added
Every output carries native 32 kHz stereo audio with lip-synced dialogue, room tone and effects, produced with the pixels rather than in a second pass. MiniMax states stable dialogue support for 11 languages, with further languages supported to varying degrees.
Two routes into a shot
The FL2VA checkpoint takes no image, one image or two: none gives text-to-video, one gives first-frame or last-frame generation, two interpolate between a first and a last frame. The Ref2VA checkpoint is the omni-reference route.
The named limits
Four to fifteen seconds per generation at 24 fps, a 768-pixel short edge by default, and 2K only by feeding the 768p result back through H3-Regenerate-2K. Only H3-Base is open-weighted; the context and regeneration stages stay hosted.
What H3 is built to do
These are the modes MiniMax documents for H3, in the vendor’s own framing.
Text to audio and video
The FL2VA checkpoint generates from text alone, with no reference material at all.
First and last frame
One image gives first-frame or last-frame generation; two images interpolate between them. In this mode the model does not add shots of its own.
Omni-reference
The Ref2VA checkpoint fuses up to nine images, three video clips and three audio clips into one prompt. Audio cannot be the sole input — it has to accompany an image or a video.
Native stereo audio
Dialogue with lip sync, room tone and effects are produced with the picture, across 11 stable languages and roughly 40 more that MiniMax says are reachable by derivation.
Editing footage you supply
MiniMax’s own reference-mode example edits an existing clip — re-animating a subject to speak a supplied line while the rest of the shot stays put. There is no separate edit tier or price for it.
Documented specifications
Every row below is stated in MiniMax’s H3 announcement or on its pay-as-you-go pricing page.
- Developer
- MiniMax
- Current model
- MiniMax H3
- Also offered
- H3 Max, text-to-video and image-to-video only
- Output duration
- 4 to 15 seconds
- Frame rate
- 24 fps
- Aspect ratios
- 21:9, 16:9, 4:3, 1:1, 3:4, 9:16
- Resolution
- 768-pixel short edge by default; 2K by regeneration
- Audio
- 32 kHz stereo, generated with the video
- Dialogue languages
- 11 stable: Arabic, Chinese, English, French, German, Italian, Japanese, Korean, Portuguese, Russian, Spanish
- Reference inputs
- Up to 9 images, 3 video clips and 3 audio clips; 12 files maximum
- Reference clip length
- 2 to 15 seconds each; video and audio total 15 seconds or less
- Open weights
- H3-Base only, under the MiniMax H3 Community Licence
Documented access channels
H3 runs as a hosted service, and one part of it is downloadable. MiniMax publishes both routes and the prices for the hosted one.
Hailuo AI
The consumer product where H3 and H3 Max are offered, alongside the image tools and the agent modes.
OpenMiniMax H3 announcement
The vendor’s own write-up of the model, covering the architecture, the three-part system and the open-source release.
OpenOpen-platform pricing
The pay-as-you-go table for MiniMax-H3, MiniMax-H3-Max and H3-Regenerate-2K, including how reference material is billed.
OpenWeights on Hugging Face
H3-Base is published as MiniMaxAI/MiniMax-H3 under the MiniMax H3 Community Licence. That licence carries territory restrictions, so read it before commercial use.
OpenThird-party platforms
H3 is also served through third-party creation platforms. The model is the same; availability, queue times and commercial terms are theirs.
MiniMax H3 questions
Is MiniMax H3 the same model as Hailuo 3.0?
Yes. MiniMax is the company, Hailuo is its consumer video brand, and H3 is the third generation of the model. The API model ID is MiniMax-H3.
Does H3 generate its own audio?
Yes. Every generation carries 32 kHz stereo audio produced with the video, including lip-synced dialogue. MiniMax states stable support for 11 languages and says further languages work to varying degrees.
How long can one H3 clip be?
Four to fifteen seconds at 24 fps. Longer sequences come from generating and extending across clips; the reference video and audio you feed in are each capped at 15 seconds in total.
What resolution does H3 actually output?
MiniMax sets the short edge to 768 pixels by default. The 2K output is produced by feeding that 768p result back through H3-Regenerate-2K rather than by generating at 2K directly.
Are H3’s weights open?
H3-Base is. MiniMax published it under the MiniMax H3 Community Licence, which is not a permissive open-source licence and carries territory restrictions. H3-Context-IR and H3-Regenerate-2K are not open-sourced and stay on the hosted side.
What does the H3 API cost?
MiniMax’s published list price is $0.08 per second at 768P and $0.13 per second at 2K; H3 Max is $0.05 at 480P and $0.08 at 768P, and the 768P to 2K regeneration is $0.05 per second. Reference audio is free and the first five reference images are free, then $0.04 each.
Can H3 work on footage I already have?
MiniMax’s reference-mode examples edit material the user supplies rather than only generating new scenes — replacing a subject, a sign or a spoken line. There is no separate editing product or price tier; treat it as part of a generation.