
What Is MiniMax H3? What It Can Make, How Good It Is, What It Costs
MiniMax H3 (Hailuo 3) generates video and audio together in one pass. What kinds of clips it makes, what it can't do, how to start, and pricing from $0.08/second.
MiniMax H3 is the AI video model MiniMax released on 31 July. You'll also see it called Hailuo 3.
The one-line difference from other AI video tools: most models hand you silent footage. H3 hands you a finished clip with sound.
Write "rain on the awning, traffic in the distance" and that's what you hear. A character speaks and the lips match. Someone walks and the footstep lands on the frame where the foot hits the ground. None of that is dubbed afterwards — picture and sound come out of the same network in one pass, instead of being stitched together from two tools.
Clips run 4 to 15 seconds, at either 768P or 2K.
What kind of clips it makes
People who talk. This is its strongest suit. A character delivers a line, the lip movement matches, and tone, room ambience and background music arrive with it. The old way took three steps — generate the footage, run AI voiceover, sync it by hand. Now it's one pass. For talking-head content, character dialogue or interview-style pieces, those are the steps you skip.
Shots with actual camera language. Push, pull, pan and track can be specified in the prompt rather than left to chance. Pair that with lighting direction (where the key light sits, what color temperature) and the output reads as cinematography instead of security footage.
Product stills turned into ads. Upload a product photo and the subject stays put while the camera moves and the light shifts. Useful for ecommerce previews and social product clips without reshooting.
The same character across multiple shots. This runs on reference material — up to 9 images, 3 videos and 3 audio files in one go (12 files max). Images lock the face and art direction, video shows how the motion should go, audio sets voice or atmosphere. For sequential storyboards and multi-shot narrative, your character stops changing faces between cuts.
One counterintuitive trap: more reference files is not better. Feed it two photos of different faces and the model oscillates between them, producing someone who resembles neither. One clean reference beats five that disagree.
What it can't do
The honest part.
No long-form. The ceiling is 15 seconds. Every long AI video you've seen online is segments generated separately and edited together. Want a minute? That's four or five generations, and you're relying on reference material to hold continuity across the seams.
No frame-accurate control. You can't specify "at 3 seconds, 12 frames, the hand reaches this position." It reads descriptions, not keyframes. Precision animation work is still out of reach for this class of model.
No reproducibility. Run the same prompt twice and you get different results. Anything you're delivering to a client still needs a human checking every clip.
Faces drift in longer takes. Fifteen seconds is the ceiling, but the longer the take, the more character consistency slips — the face gradually stops looking like the one you started with. Always check this on longer shots.
How to start using it
Fastest path is a browser. Open a page and generate — no environment setup, no GPU, no model download. This site (minimax3.com) is one of those entry points, with free credits on signup.
To be clear: this is not the Hailuo AI official site, it's an independent third-party tool. For the official entry point go to MiniMax's own domains; credits here don't transfer to an official account either.
If you're building it into a product, use the official API with model name MiniMax-H3. It's asynchronous — submit a task, poll for the result.
The third path is local deployment, and that one isn't open to everyone.
Before you plan on local deployment, check your country
On 3 August MiniMax put the H3 weights on Hugging Face — 33 billion parameters, with ComfyUI support landing the same day. So the "when will H3 be open sourced" and "open weights countdown" posts you can still find are out of date.
But the release comes with a condition. The license is the MiniMax H3 Community License: free for non-commercial use, and commercial use is allowed for companies under $20 million in annual revenue. Permissive enough — except it defines an "Applicable Territory," and the United States, the European Union, the United Kingdom and South Korea are not in it.
Not being in it means users in those four places aren't licensed to run the weights locally, and can't distribute what they produce locally either. The license doesn't state why those regions are excluded.
So if that's you, the ComfyUI deployment tutorials you find don't apply. The browser route or the official API are the paths that do (both run on service terms, which are separate from the weights license). If you're seriously planning to deploy, read the license text yourself — this isn't legal advice.
Outside those regions, local works fine. ComfyUI needs 0.30.0 or later, and there are three official workflow templates. The smallest usable file set is 42.5 GB (123.6 GB at full precision), and ComfyUI reports 12 GB of VRAM with offloading is enough to run it.
What it costs
The official API bills per second:
| Quality | Price |
|---|---|
| 768P | $0.08 / second |
| 2K | $0.13 / second |
A 10-second 2K clip is about $1.30; a 5-second 768P clip about $0.40. That low cost of iteration matters more than it sounds — you can tune the prompt at 768P and only move to 2K for the final.
Running locally costs no API fees, but you need the 42.5 GB and a card that can handle it.
Free credits on third-party web tools — this site included — run on their own ledger and don't touch your MiniMax account.
How to write prompts that actually work
This step decides output quality, and most people get it wrong at first.
A typical bad prompt:
A woman on a rainy city street, cinematic, 4K, hyperrealistic
Looks like plenty of information, but there's nothing to execute. What shot is "cinematic"? "4K" is a resolution parameter that doesn't belong in a prompt. Write it this way and the model is guessing.
Same scene, written properly:
A woman in a beige trench coat shelters under a convenience store sign, closing her umbrella with her right hand. Night, a Tokyo back alley, standing water reflecting neon. Medium shot, camera pushes slowly to a waist-up framing. Key light from the sign behind her, cool white, rain visible in the backlight. Sound: rain on the awning, traffic in the distance. Ends as she looks up toward something off-camera.
Every clause maps to something specific in the frame. Think in this order:
- Who or what — specific enough to recognize
- What they're doing — an action with a start and an end
- Where — time of day, weather
- Camera — shot size, how it moves
- Light — direction, color temperature
- Sound — dialogue, ambience
- How it ends — where the last frame lands, so it can cut to your next shot
The prompt ceiling is 7,000 characters, but filling it doesn't help — long prompts bury the priority. Get one shot right first.
Once you have a clip, check four things: whether the person stays the same throughout, whether the action completes, whether lips match the dialogue, and whether the final frame can cut to the next shot. When one fails, change only that line — rewriting the whole prompt changes the parts that already worked.
Only upload reference material you have rights to. Real people's photos, other people's work, trademarked products — clear those first.
Things people mix up
H3 is not MiniMax M3. M3 is the same company's language model, for chat and code. Nothing to do with video. The names are close enough that searches collide constantly.
It's not Tencent Hy3 either. Different model, different company. Some API catalogs list them next to each other.
The 1440p figure floating around is wrong. Official tiers are 768P and 2K. The 1440p number most likely came from an early miscalculation of 2K.
Is it worth trying
If you're making 15-second clips that need picture and sound to line up, H3 is currently the low-effort option — it removes the separate voiceover and sync steps.
If you need multi-minute output, frame-level control, or deterministic results you can hand to a client, no current AI video model covers that, H3 included.
To evaluate it quickly: write one layered prompt like the example above, generate 5 seconds at 768P for about 40 cents, run the four checks, then change exactly one thing and generate again. Both runs together cost under a dollar — enough to tell whether it fits your work.
This site runs H3 in the browser with free credits on signup. For API integration, see the MiniMax H3 API guide. Specs can be verified against the official API docs and the Hugging Face model card.
minimax3.com is an independent third-party tool with no affiliation with, endorsement by, or operational relationship to MiniMax or Hailuo AI. Product names are used descriptively; copyright belongs to their respective owners.
Generate your first image with MiniMax H3 — right now
Reliable non-Latin text rendering, directed editing, and 50+ ready-to-use prompts. No downloads — just open in your browser.