MiniMax H3 Max
MiniMax H3 Max renders a five-second 768P clip with synchronized audio in under three seconds. Generate with it below, see what it makes, and check what it costs before you build anything on it.
MiniMax H3 Max output, 1344×768 with sound — unmute to hear it.
Run MiniMax H3 Max
Pinned to this model — 480P or 768P, 5 to 15 seconds, audio included. Draft at 480P while the prompt is still moving; it is the tier where this model is genuinely cheaper.
Video Generation Settings
Model
0/7000
Video Resolution
Duration
Video Aspect Ratio
Generate Audio
AI will generate matching audio for the video.
Video Preview
Click video to view full size • Swipe or use arrows to see more
MiniMax H3 Max capability guide
What this model accepts and produces. Everything here is a hard limit of the model, not a setting on this page.
Output
- 480P or 768P — there is no 2K tier at any price
- 5 to 15 seconds; 5 is the floor, not 4
- 1344×768 at 24fps on the 16:9 setting
- Synchronized audio on every generation, with no toggle
Inputs
- Text to video takes all six aspect ratios: 21:9, 16:9, 4:3, 1:1, 3:4, 9:16
- Image to video takes a first frame, and optionally an end frame
- Image to video inherits the aspect ratio of the image you pass in — crop before uploading
- No published prompt-length limit; treat a very long brief as untested
Supported combinations
- Text only
- Text + first frame
- Text + first frame + end frame
The generator above prices every combination in credits before you spend anything.
MiniMax H3 Max capabilities
Every clip below is model output at its native 1344×768, each with its own audio track.
Character consistency across a shot
A stylised character holds the same face, hair and wardrobe as the camera tracks with them. Drift between the first and last second is the usual failure mode here, and it is what the image-to-video leaderboards measure.
First and last frame
Pin where a shot starts and where it ends and the model fills the middle. Physical processes with a clear start and end state — something inflating, opening, assembling — are where pinning both ends beats describing the motion in words.
Audio in the same pass
Every generation returns synchronized sound — room tone, foley, ambience and music — rather than a silent clip you score afterwards. Unmute the player to hear it. There is no toggle and no separate audio job.
Holding an art direction
A commercial-looking spot keeps one lighting and grading treatment from beginning to end instead of drifting between looks mid-clip. This is the property that decides whether several generated shots can be cut together.
Text rendered into the frame
Short on-screen type can come out clean and correctly spaced, which is unusual for a video model. Keep it to a few words — the next section shows where this stops working.
Where the text rendering falls over
The same feature that produced clean type in the previous clip produces "the takes hells war" here — words in roughly the right shape, arranged into something that is not a sentence. Longer strings, serif faces and multi-line layouts all push it this way. Treat in-frame text as a couple of words at most, and read every frame before shipping a clip that depends on it.
MiniMax H3 Max examples
MiniMax H3 Max vs MiniMax H3
It was post-trained from H3’s open weights, so the two are close relatives.
| MiniMax H3 Max | MiniMax H3 | |
|---|---|---|
| Lineage | Post-trained from H3’s open weights | MiniMax |
| Duration | 5-15 seconds | 4-15 seconds |
| Max resolution | 768P | 2K |
| Low resolution | 480P | 768P |
| 5s at 768P, wall time | Under 3 seconds | Roughly 90-120 seconds |
| Native audio | Yes, synchronized | Yes, dual-channel stereo |
| Aspect ratios | Six on text-to-video; image-to-video inherits from the input image | Six on text-to-video; image-to-video inherits from the input image |
| Reference inputs | First and last frame; reference-to-video added days later | Up to 12 files: 9 images + 3 videos + 3 audio |
| Prompt length | Not published | ~7,000 characters |
| Design Arena i2v Elo | 1,341 | 1,333 |
| Continuous mode | Director, up to 2 minutes | None |
The full comparison, including what twenty attempts cost and take on each, is in H3 Max vs MiniMax H3. For 2K delivery you want MiniMax H3 itself.
Settings and steering
prompt_expansion_modebalanced | qualityBalanced adds about a second of overhead; quality adds about thirty — ten times the render itself. Steer on balanced and switch to quality only for the take you intend to keep. Left on quality it throws away the reason to come here at all.
resolution480P | 768PThere is no 2K. 480P is the drafting tier and the cheapest way to iterate; 768P costs roughly half again as many credits per clip. Find the shot at 480P, confirm it at 768P.
aspect_ratio21:9, 16:9, 4:3, 1:1, 3:4, 9:16Text-to-video only — on both models. Image-to-video has no such parameter on either and follows the ratio of the image you pass in, so vertical output needs a vertical first frame.
duration5-15s, default 5Five is the floor. Shorter beats mean generating five seconds and trimming. Nothing lifts the 15-second ceiling, so longer pieces are still stitched.
end_image_urlimage-to-video onlyPins where the shot ends. Pair it with image_url to lock both ends of a movement instead of describing it and hoping.
Change one clause per run
On a slow model you front-load the prompt because a round trip is expensive. Here the opposite pays: alter a single clause, look, alter the next. When two things change at once and the shot improves, you have learned nothing about which one did it.
Fix the subject before you touch the camera
Camera language is the most volatile part of a video prompt — it re-rolls composition, framing and pacing together. Get the subject and setting stable first, then introduce the move.
Draft at 480P, confirm at 768P
The two tiers do not always agree on fine detail, so a prompt that lands at 480P is a candidate rather than a finished take. Budget one 768P confirmation run per shot you keep, and re-check faces and any in-frame text at the higher tier.
Lock the seed once you are comparing
A random seed while you are still finding the shot gives you variety for free. Once you are comparing two phrasings, fix the seed so the difference you are reading is the prompt and not the roll.
The prompt structure is unchanged from the parent model — our H3 prompt guide covers the format both accept, the first and last frame builder writes the keyframe form, and the camera control reference lists all 20 moves H3 recognises.
How to use MiniMax H3 Max
- 01Write the shot, not the sceneOne subject, one action, one camera move. The model tops out at 15 seconds, so a brief describing three things happening in sequence gets compressed rather than filmed.
- 02Draft at 480PGenerate at the cheaper tier while the prompt is still moving. At roughly three seconds a run you can try ten phrasings in the time one slow generation used to take.
- 03Change one clause at a timeAdjust a single element per run so you can tell which change did the work. This is the technique a fast loop rewards and a slow one punishes.
- 04Confirm the keeper at 768PRe-run the winning prompt at the higher tier and check faces, hands and any in-frame text before you treat it as final.
Mistakes that cost money
Leaving prompt expansion on quality while iterating
Turns a three-second loop into a thirty-second one — the fastest way to lose the entire speed advantage.
Iterating at 768P
Every re-roll costs about half again as many credits as the same shot at 480P. Draft low, confirm high.
Treating Director like a clip endpoint
It bills a whole session with a minimum length, not the seconds anyone watched — budget it by the minute before you design around it.
Uploading a landscape still and expecting vertical output
Image to video inherits the aspect ratio of the input. Crop first, or the whole run is wasted.
Shipping a clip whose on-screen text you did not read
Text rendering degrades fast past a couple of words, and a garbled line is invisible until someone reads it.
Planning a deliverable at 2K
There is no 2K tier at any price. Discovering that after the shot is approved means re-rendering it elsewhere.
MiniMax H3 Max questions
- What is MiniMax H3 Max?
- A video model post-trained from MiniMax H3’s open weights and released on 1 September 2026. It renders a five-second 768P clip with synchronized audio in under three seconds. "Max" does not mean it is the larger model in the lineup — it trades 2K away for speed.
- How is MiniMax H3 Max different from MiniMax H3?
- Speed and resolution. It returns the same clip far faster, but caps at 768P where MiniMax H3 reaches 2K. Its 480P tier is the cheapest way to iterate on this site; 2K work still belongs on MiniMax H3.
- What resolution and duration does MiniMax H3 Max support?
- 480P or 768P, 5 to 15 seconds, with synchronized audio on every generation. At 16:9 and 768P the output is 1344x768 at 24fps. There is no 2K tier.
- How much does MiniMax H3 Max cost here?
- In credits, and the generator prices every combination before you spend anything — the button shows the exact cost for the resolution and duration you picked. A five-second 480P clip is the cheapest thing on the page; longer clips and 768P scale up from there.
- Can I generate with MiniMax H3 Max here?
- Yes. The generator on this page is pinned to MiniMax H3 Max, and new accounts get free credits on signup. A five-second 480P clip is the cheapest way to see what the model does with your prompt.
- Which MiniMax H3 Max endpoints exist?
- Text to video and image to video came first, with reference to video following days later. Director is the continuous mode — it keeps context across a streaming session rather than returning a file.
- What is MiniMax H3 Max Director?
- The continuous version. It holds one video stream open and maintains context across it, so prompts sent mid-stream change the action without breaking continuity. Sessions cap at two minutes as standard, and it is metered by the session rather than by the clip.
- Is MiniMax H3 Max really ranked first?
- On two boards, narrowly. Design Arena puts it first for image-to-video at 1,341 Elo with MiniMax H3 right behind at 1,333, and Artificial Analysis ranks it first for image-to-video with audio at 1,201 Elo across 2,177 samples. Eight Elo over its own parent is a lead, not a gulf — the speed is the real story.
- What settings give the best results with MiniMax H3 Max?
- Keep prompt_expansion_mode on balanced while you iterate, because quality adds about thirty seconds per run. Draft at 480P and confirm the shot you keep at 768P. On image-to-video, crop the first frame to the aspect ratio you want, since the output inherits it.
- Can MiniMax H3 Max render text in the frame?
- Sometimes. A couple of short words can come out clean, but longer strings degrade into letter-shaped nonsense — there is an example of exactly that on this page. Read every frame before shipping a clip whose meaning depends on the text.
This site is an independent third-party tool, not affiliated with MiniMax.