MiniMax H3Long Video Generator
H3 renders 4 to 15 seconds per generation. Everything longer is segments joined end to end. This planner does the split, writes the right prompt for each segment — including the continuation line almost everyone gets wrong — and tells you what it will cost before you spend anything.
The short answer
Can MiniMax H3 make a 60-second video?
Not in a single generation. The hosted H3 endpoint accepts 4 to 15 seconds, whole seconds only, and that is the hard ceiling. A 60- or 120-second H3 video is always several clips joined end to end, where each clip opens on the final frame of the one before it. That technique is called video continuation, and it works — but it is a chain of separate inferences, not one long render.
- One generation: 4–15 seconds, with native synchronised audio.
- 60 seconds = 4 segments at 15s. 120 seconds = 8 segments at 15s.
- Each segment after the first is an image-to-video job whose input image is the previous segment’s last frame.
This page plans the chain and writes the prompts. Running the segments and joining the files is still something you do yourself, one generation at a time — we would rather say so than sell you a button that does not exist yet.
The planner
Split a runtime into segments H3 will actually accept
Set a target length, describe who and where, then give each segment one beat. The prompts below are ready to paste — English, because that is the format H3 was trained on.
One beat per segment
What changes during this segment. Write the action, not the summary — a segment with nothing to do will drift or loop.
Your segments
4 segments · 60s · about 804 credits total
Run them in order. After each one finishes, take its final frame and use it as the input image of the next.
Segment 1 · 0s–15s · 15s
For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced. integrated_multimodal_description: [Shot 1] Live-action, cinematic, a courier in a red rain jacket with a frayed left cuff, mid-twenties, short black hair, in a neon-lit night market street, wet asphalt, stalls still open, matching the position, framing, and lighting established by <Picture 1>. she pushes off the kerb and threads between two stalls, checking the address on her phone. overall_soundscape: Steady rain on canvas awnings, tyres through standing water, market chatter fading in and out, a tram bell in the distance. non_diegetic_music: N/A
Segment 2 · 15s–30s · 15s
For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced. integrated_multimodal_description: [Shot 1] Live-action, cinematic, a courier in a red rain jacket with a frayed left cuff, mid-twenties, short black hair, in a neon-lit night market street, wet asphalt, stalls still open, continuing without interruption from the position, framing, motion direction, and lighting established by <Picture 1>. she slows at a crossing as a tram passes, rain beading on the jacket. The identity, clothing, colour palette, key props, and lens character stay identical to <Picture 1>; no cut, no restart, no change of shot scale at the opening. overall_soundscape: Steady rain on canvas awnings, tyres through standing water, market chatter fading in and out, a tram bell in the distance. non_diegetic_music: N/A
Segment 3 · 30s–45s · 15s
For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced. integrated_multimodal_description: [Shot 1] Live-action, cinematic, a courier in a red rain jacket with a frayed left cuff, mid-twenties, short black hair, in a neon-lit night market street, wet asphalt, stalls still open, continuing without interruption from the position, framing, motion direction, and lighting established by <Picture 1>. she lifts the parcel from the crate and scans the shopfront numbers. The identity, clothing, colour palette, key props, and lens character stay identical to <Picture 1>; no cut, no restart, no change of shot scale at the opening. overall_soundscape: Steady rain on canvas awnings, tyres through standing water, market chatter fading in and out, a tram bell in the distance. non_diegetic_music: N/A
Segment 4 · 45s–60s · 15s
For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced. integrated_multimodal_description: [Shot 1] Live-action, cinematic, a courier in a red rain jacket with a frayed left cuff, mid-twenties, short black hair, in a neon-lit night market street, wet asphalt, stalls still open, continuing without interruption from the position, framing, motion direction, and lighting established by <Picture 1>. she hands the parcel over at a lit doorway and steps back into the rain. The identity, clothing, colour palette, key props, and lens character stay identical to <Picture 1>; no cut, no restart, no change of shot scale at the opening. overall_soundscape: Steady rain on canvas awnings, tyres through standing water, market chatter fading in and out, a tram bell in the distance. non_diegetic_music: N/A
Opens the image-to-video generator with this prompt already loaded. Sign-in and credits are needed to generate.
Before you spend credits
Nothing to flag — this plan is as continuous as a chain gets.
The mechanic
What “video continuation” actually does
There is no extend parameter on the H3 API. Continuation is a convention, not an endpoint: you hand the next generation a single still — the last frame of the previous clip — and ask it to carry on.
- 1
Generate segment 1
Text-to-video, or image-to-video from your reference still. This segment sets the identity, palette, and lens for everything that follows.
- 2
Take its final frame
Pause the finished clip at the very end and export that frame, or pull it in any editor. This still becomes the next segment’s input image.
- 3
Run the next segment as I2VA
Image-to-video, with the final frame as the input image and the continuation prompt from the planner. Not first-and-last-frame — there is no second picture yet.
- 4
Repeat, then join the files
Loop until you reach the target runtime, then concatenate the clips in any editor. The segments are separate MP4s until you join them.
What carries across a seam
- Composition and framing at the cut — it is literally the same frame.
- Anything you repeat verbatim in the prompt: subject, clothing, location, style token.
- Broad colour and lighting, as long as the frame you hand over is well exposed.
What resets at every seam
- Motion. A still carries no velocity, so the model re-infers the movement and fast action can visibly stall at the cut.
- Audio. Each segment generates its own; ambience and any music restart unless you describe them identically.
- Everything off-screen. Anything not visible in the handed-over frame no longer exists as far as the next segment is concerned.
This is the rule most long-video guides get wrong: only segment 1 — or a segment that has to land on a supplied end frame — is FL2VA. Every continuation segment is I2VA, and I2VA uses a different instruction line. Use the FL2VA sentence with one picture and the frame gets quietly ignored.
Expectations
How far a chain holds together
Each segment sees exactly one frame of history and your prompt — nothing else. That single fact predicts what breaks, and roughly when.
| Chain depth | Usually still holds | Starts to go |
|---|---|---|
| Segments 1–2 | Identity, wardrobe, palette, location. | Motion pace across the first seam, if the action was fast. |
| Segments 3–4 | Identity and location, given a repeated anchor. | Colour temperature drifts; secondary props change shape. |
| Segments 5–6 | The broad scene and the described anchor. | Faces soften and reshape; clothing details rewrite themselves. |
| Segments 7–8 | Style token and general setting. | Cumulative drift is usually obvious side by side with segment 1. |
These are the failure modes the mechanism predicts and that show up in community reports — not measured rates from a benchmark we ran. When we publish numbers, they will say how they were produced.
Cost
What a long H3 video costs here
H3 bills by generated second, so a chain costs the same as its total runtime — plus every segment you regenerate. Regenerations are the real budget line, not the plan.
| Runtime | Segments at 15s | Credits | If you redo two links |
|---|---|---|---|
| 30s | 2 | 402 | 804 |
| 60s | 4 | 804 | 1206 |
| 90s | 6 | 1206 | 1608 |
| 120s | 8 | 1608 | 2010 |
Figures come from this site’s live H3 pricing, so they match what the generator will actually charge. A failed generation is refunded automatically; one you simply dislike is not.
How to
Build a 60-second H3 video
- 1
Plan the chain
Choose 60 seconds and 15-second segments to get four links, then write one beat per segment so no segment has empty time to fill.
- 2
Lock the identity anchor
Describe the subject once, concretely, and let the planner repeat it verbatim in all four prompts. This is the single highest-leverage thing you can do about drift.
- 3
Generate segment 1 from a reference image
Open the image-to-video generator, attach your start frame, paste the segment 1 prompt, and generate at 15 seconds.
- 4
Chain the rest
Export the final frame of each finished clip, use it as the input image for the next segment, and paste that segment’s prompt. Repeat until the runtime is covered.
- 5
Join the clips
Concatenate the four MP4s in any editor. Trim a few frames at each seam if the motion restart is visible.
Be precise
Three different things people call “long H3 video”
These get mixed together constantly, and only one of them is something you can do today.
| Approach | What it means | Where it stands |
|---|---|---|
| Single generation | One inference, one file, native audio throughout. | Available — capped at 15 seconds. |
| Chained segments | Several generations, each opening on the previous final frame. | Available today. This page plans it; you run and join the parts. |
| Streaming / causal generation | Video produced continuously as it plays, with no fixed clip length. | Research-stage. Not available in the hosted H3 API, and not offered here. |
FAQ
MiniMax H3 long video questions
What is the maximum length of a MiniMax H3 video?
Fifteen seconds per generation, four seconds minimum, whole seconds only. That applies to text-to-video, image-to-video, and reference-to-video alike. Anything longer is segments joined together.
How do I make a 2-minute MiniMax H3 video?
Plan it as eight 15-second segments. Generate the first from a reference image, then use each finished clip’s final frame as the input image of the next, with a prompt that repeats the same identity anchor. Join the eight files at the end. Expect to regenerate one or two links.
Does H3 have an extend or continue endpoint?
No. There is no extend parameter on the hosted API. Continuation is done by handing the next generation the previous clip’s last frame as its input image — which is why the continuation prompt is an image-to-video prompt, not a first-and-last-frame one.
Can I upload my own video to continue it?
Not on this site right now. The generator takes images, not video. If you have a clip you want to extend, export its final frame and start the chain from that image instead.
Why does the motion stutter at the joins?
Because a still frame carries no velocity. The next segment re-infers the movement from scratch, so fast action can appear to pause at the cut. Keeping the pace moderate across seams, and trimming a few frames when you join, both help.
Does the audio stay continuous?
Not by itself. Each segment generates its own audio, so ambience restarts at every seam unless the same soundscape line appears in every prompt — which is why the planner repeats it for you. For music, scoring the joined video in an editor is more reliable.
Is this an official MiniMax feature?
No. The 4–15 second limit is MiniMax’s. Segmented continuation is a community workflow built on top of it, and this planner is our implementation of it. This site is not affiliated with MiniMax.
Keep reading
More on H3 prompts and frames
First & last frame prompt builder
The FL2VA counterpart to this page: when you have both ends of a single shot and need the middle written correctly.
The full H3 prompt format
The three-field structure shared by all four modes, plus the camera and dialogue vocabulary H3 was trained to read.
First and last frame, end to end
How to choose the two frames, six use cases that actually pay, and what to do when the last frame refuses to land.
Turbo LoRA and step counts
Running H3 locally in 4–8 steps instead of ~20, and the point where the distilled checkpoint stops being a win.
Plan it here. Generate it next door.
The generator on this site runs H3 image-to-video with a start frame and an optional end frame — exactly what a continuation segment needs. Plan the chain above, then run the segments one at a time.
The prompt grammar on this page — the I2VA instruction line, the FL2VA alignment sentence, two-decimal durations, the three output fields — comes from MiniMaxAI’s Video Prompt Writing Guide shipped with the H3 open weights, checked as of August 2026. The 4–15 second limit is the hosted API’s. This site is not affiliated with MiniMax.