
MiniMax H3 Prompt Guide: 6 Workflows for Better Videos
How to write prompts for MiniMax H3 video generation — three input modes, a reusable prompt formula, and six production workflows with example prompts.
Most failed H3 generations are not model failures. They are prompt failures. The model did exactly what was asked — the problem was the ask.
MiniMax published an official prompting guide on 3 August 2026, alongside the open weights release. This article distills that guide into a decision framework, a reusable prompt template, and six production workflows with example prompts you can adapt immediately. If you have not set up API access yet, start with the MiniMax H3 API guide.
Choosing your input mode
H3 accepts three kinds of input. Each one changes what the prompt needs to specify and what it can leave out. Picking the wrong mode is a common source of wasted generations.
Text-to-video
No images, no video, no audio. The prompt carries the entire creative brief.
Use this when you are working from imagination or written scripts and have no visual reference to upload. Because the model has nothing to look at, every visual detail matters — subject appearance, setting, action, camera movement, lighting, sound design.
Text-to-video requires you to specify an aspect ratio explicitly (21:9, 16:9, 4:3, 1:1, 3:4, or 9:16). The model cannot infer it from a missing image.
First-frame / last-frame (image-to-video)
Supply a start image, an end image, or both. The prompt explains what happens between them.
Use this for product demos where the product must look exactly right, character state changes (standing to sitting, day to night), scene transitions, or animating a still poster. The image locks the visual that matters most, and the prompt fills in the motion and sound.
The output aspect ratio follows the input image — you do not set it separately. First-frame and last-frame inputs cannot be combined with reference inputs in the same request.
Omni-reference (reference-to-video)
The most flexible mode. Upload a mix of images (up to 9), videos (up to 3, combined 15 seconds max), and audio files (up to 3, combined 15 seconds max). Then use the prompt to assign each upload a role.
Use this when you have specific assets to preserve — a character's face from a photo, a camera movement from a video clip, a voice or ambient track from an audio file. The prompt becomes a casting sheet: it tells the model which input provides which element.
Audio cannot be the only reference; at least one image or video must be included.
When to use which
If you have no visual assets and are working from a script or concept, use text-to-video.
If you have one or two key frames that must appear exactly as-is in the output — a product photo, a character portrait, a background plate — use first/last frame.
If you have multiple reference assets (character photos, motion reference clips, voice recordings, style images) and need the model to combine them, use omni-reference.
A rule of thumb: the more reference material you supply, the less visual description the prompt needs — but the more role-assignment language it needs. Text-to-video prompts are long on description. Omni-reference prompts are long on instructions.
The prompt formula
MiniMax's official guide breaks a complete prompt into three parts:
Complete prompt = Reference material description + Core creative idea + Visual process description
This formula works across all three input modes, with different emphasis for each.
Part 1: Reference material description
When you upload files, number and label them. The model needs to know what each file is for.
Pattern: @image 1 is the character reference — preserve face, hair, and clothing. @video 1 provides the camera movement — follow its push-in timing. @audio 1 is the ambient track — match its rhythm.
Call out the specific features to preserve. "Use this image as reference" is too vague. "Use this image for the character's face and outfit, but not the background" gives the model a constraint it can follow.
For text-to-video, skip this section entirely. For first/last frame, specify which image is the start frame and which is the end frame.
Part 2: Core creative idea
One sentence that covers subject, scene, event, style, and camera approach. This is the thesis of the video.
Example: A woman in a tailored navy coat walks through a rain-soaked alley at dusk, shot in a slow tracking medium from behind, with warm tungsten light from a doorway ahead.
If the core idea takes more than two sentences, you are probably describing multiple shots. Split them.
Part 3: Visual process description
A time-stamped or shot-by-shot breakdown of what happens frame by frame.
Example:
- 0–3s: Close-up of hands adjusting a coat collar. Camera holds still.
- 3–6s: Pull back to medium shot. Character turns and walks toward the light. Camera tracks slowly.
- 6–10s: Cut to the doorway ahead. Warm light spills across wet pavement. Character enters frame from the left.
This section is where most prompts fall short. People write a mood statement ("cinematic and emotional") instead of a shot plan. The model executes actions, not adjectives.
Sound control goes here too. You can specify background music genre, ambient sounds, spoken dialogue, or suppress music entirely with "non-narrative music: N/A."
The copy-paste template
[REFERENCE DESCRIPTION]
@image 1: [role — what to preserve]
@image 2: [role — what to preserve]
@video 1: [role — what to extract]
@audio 1: [role — what to match]
[CORE IDEA]
[Subject] [action] in [setting], shot as [camera style], with [lighting/mood].
[VISUAL PROCESS]
0–[X]s: [Shot type]. [Action]. [Camera movement]. [Sound].
[X]–[Y]s: [Shot type]. [Action]. [Camera movement]. [Sound].
[Y]–[Z]s: [Shot type]. [Action]. [Camera movement]. [Sound].
[SOUND DESIGN]
BGM: [genre/mood or "N/A"]
Ambient: [specific sounds]
Dialogue: [lines, if any]Delete any section that does not apply to your input mode. For text-to-video, remove the reference description block and compensate with more visual detail. For first/last frame, replace the reference block with frame assignments.
Six production workflows
The following workflows come from MiniMax's official guide, expanded with prompt patterns you can modify for your own projects. Each targets a different production need.
1. Brand films and cinematic expression
The goal is camera language — motivated cuts, deliberate transitions, visual continuity across shots.
This workflow relies heavily on first/last frame mode. Keyframes lock the start and end of each shot, and the prompt fills in the camera movement and transitions between them. Hard cuts, blackouts, flash transitions, and motivated camera pushes all live in the prompt's visual process section.
Useful camera vocabulary for H3 prompts: push-in, pull-back, slow pan left, tracking shot, rack focus, Dutch angle, bird's eye, worm's eye, Steadicam follow, whip pan, static locked-off.
Example prompt (text-to-video):
A man in a dark wool overcoat stands at the edge of a rooftop at
golden hour. City skyline behind him, backlit.
0–3s: Extreme close-up of his eyes. Shallow depth of field. The
city lights reflect in his pupils. Camera holds still. Sound: wind,
distant traffic.
3–5s: Black screen, 0.5s. Then cut to wide shot from behind —
the man silhouetted against the skyline. Camera slowly pushes in.
5–8s: Medium shot from the side. He turns his head to the right.
Key light from the setting sun, warm amber. Camera tracks around
to face him.
8–10s: Close-up of his hand releasing a paper airplane. Camera
follows it off the rooftop into the sky. Sound: paper flutter, then
silence.
BGM: minimal piano, single repeating note
Ambient: rooftop wind, faint city hum2. Visual creative and content packaging
Stylized treatments — retro anime, hard-edge silhouettes, comic collage, asymmetric split screens. The prompt drives style through specific visual language rather than style tags.
This workflow benefits from omni-reference mode. Upload a style reference image and a motion reference video, then let the prompt assign roles.
Example prompt (omni-reference):
@image 1: style reference — preserve the flat color palette, thick
ink outlines, and halftone dot texture.
@video 1: motion reference — match the timing of its zoom-outs
and the rhythmic cuts on every beat.
A street musician plays saxophone under a bridge at night, drawn
in retro anime style with visible ink grain.
0–3s: Asymmetric split screen — left panel shows hands on keys
in close-up, right panel shows the full figure from a low angle.
3–5s: Panels merge into single frame. Camera pulls back to
reveal the bridge and the river below. Hard-edge silhouette
transition.
5–8s: Comic-panel collage — four rapid cuts of saxophone
close-ups, each held for 0.5s, framed with thick black borders.
8–10s: Return to single wide shot. Text "MIDNIGHT SET" appears
bottom-center in hand-drawn lettering. Hold 2s.
BGM: jazz saxophone, synced to the visual rhythm
Ambient: river water, echo under the bridge3. Brand and fashion content
The product enters the scene through the character's action, not through a static display. The camera follows the character; the product is revealed because the character uses it, reaches for it, or wears it.
Upload the product as an image reference and specify its role precisely. The prompt should describe how the product appears naturally in the scene.
Example prompt (omni-reference):
@image 1: character reference — preserve face, hairstyle, and
body proportions.
@image 2: product reference — this is a leather crossbody bag in
cognac brown. Preserve the color, buckle detail, and strap width.
A woman walks along a desert highway at sunset, wearing a linen
shirt and wide-leg trousers. She carries the bag from @image 2 on
her left shoulder.
0–3s: Wide shot from the front. She walks toward camera. A
vintage car is parked on the roadside behind her. Golden hour
light, long shadows.
3–6s: Medium shot from the side. Her hand adjusts the bag strap.
Camera pans to follow. The bag fills the lower third of the frame.
6–9s: Close-up of the bag — the buckle catches sunlight. Shallow
depth of field. Camera holds, slight tilt up to her face. She
glances down and smiles.
9–10s: Pull back to wide. She continues walking. Dust rises from
her steps.
BGM: acoustic guitar, unhurried
Ambient: desert wind, gravel footsteps
Dialogue: none4. Animation and stylized imagery
Character reference images lock face, hair, body, and clothing. The model maintains these across the entire clip. Suitable for character PVs, game cinematics, anime PVs, and IP content where visual identity must not drift.
The key prompt technique is explicit feature enumeration: list every attribute the model must preserve rather than saying "keep the character consistent."
Example prompt (omni-reference):
@image 1: character reference — preserve exactly: silver-white
bob-cut hair, violet eyes, sharp jawline, black turtleneck, silver
pendant necklace.
A character with @image 1's appearance stands in a neon-lit
cyberpunk alley. Rain falls. Reflections on wet ground.
0–4s: Medium shot, frontal. She looks directly at camera. Rain
hits her shoulders. Neon signs reflect on her face — alternating
pink and blue. Camera slowly pushes in.
4–7s: Cut to profile view. She raises her right hand. A
holographic interface appears from her palm. Camera tracks from
her face to the hologram.
7–10s: Over-the-shoulder shot looking at the hologram. Data
scrolls across it. Camera slowly rotates around to face her
through the translucent display.
BGM: ambient synth, low and pulsing
Ambient: rain, neon buzz, distant sirens5. Product and e-commerce marketing
Start with product photos. The prompt turns static images into 360-degree displays, material close-ups, ergonomic demonstrations, and real-environment placements.
This workflow uses first/last frame mode or omni-reference depending on how much camera control you need. A single product photo as the first frame works for straightforward turntables; multiple product angles as references work for complex reveals.
Example prompt (first-frame):
First frame: product photo — a wireless ergonomic mouse on a
white surface, viewed from a 45-degree angle.
The mouse sits on a clean white desk. The camera showcases the
product from multiple angles.
0–3s: Camera slowly orbits the mouse clockwise, 90 degrees.
Clean white background. Soft diffused light from above.
3–5s: Close-up of the side grip texture. Camera pushes in until
the surface pattern fills the frame. Sound: soft click of the side
button.
5–8s: Pull back. A hand enters from the right and naturally grips
the mouse. Medium shot showing the ergonomic curve against the
palm. The desk surface changes to dark walnut.
8–10s: Wide shot of a full desk setup — monitor, keyboard, and
the mouse in use. The hand clicks and scrolls. Warm office
lighting, shallow depth of field on the mouse.
BGM: N/A
Ambient: quiet office, keyboard tapping, mouse click6. Character, object, and scene editing
Green screen removal, background replacement, dialogue replacement. The prompt specifies what stays and what changes.
For background replacement, match the character's existing lighting in the prompt — if the character is lit from the left, the new background should have its light source on the left. Mismatched lighting is the most common tell.
For dialogue replacement, upload the target audio as a reference. The prompt asks the model to adjust the character's lip movement and performance to match the new audio.
Example prompt (omni-reference, background replacement):
@image 1: character on green screen — preserve the character's
appearance, pose, and left-side key lighting exactly.
@image 2: new background — a Japanese garden in autumn, warm
afternoon light from the left.
Place the character from @image 1 into the scene from @image 2.
Match the lighting direction — key light from the left in both.
Remove the green screen entirely.
0–5s: Medium shot. Character stands in the garden. Maple leaves
drift down around her. Gentle camera sway, handheld feel. Sound:
birdsong, light wind through leaves.
5–10s: Slow push-in to close-up. Leaves continue falling. One
leaf lands on her shoulder. She brushes it off. The lighting on
her face matches the garden's warm afternoon tones.
BGM: traditional koto, very quiet
Ambient: garden birds, wind, rustling leaves
Dialogue: noneCommon mistakes
Mood words instead of actions. "Cinematic and emotional" gives the model nothing to execute. "Camera pushes in slowly over 3 seconds while the subject lowers their gaze" does.
Unlabeled references. Uploading three images without specifying which is the character, which is the style, and which is the background. The model picks arbitrarily.
Conflicting references. Two face photos of different people. The model averages them, and the result looks like neither.
Filling the prompt with resolution or quality tags. Terms like "4K," "hyperrealistic," or "8K ultra HD" do not map to model parameters. They waste prompt space.
Overloading a single generation. A 10-second clip with five scene changes, three characters, and a dialogue exchange. Simplify to one clear dramatic beat per generation. Edit the clips together afterward.
Ignoring the ending frame. If you plan to cut this clip into a sequence, specify where the last frame lands. An unplanned ending makes editing harder.
Further reading
- MiniMax H3 API guide — endpoint setup, parameters, and async polling
- MiniMax H3 open weights — downloading from Hugging Face, ComfyUI setup, license restrictions
- What is MiniMax H3 — capabilities, limitations, and pricing overview
- Official API documentation
- Hailuo AI — browser-based generation
- MiniMax Hub — community prompts and examples
The prompt template above covers most production scenarios. Start with one shot, verify the output against the four checks — identity, action completion, audio sync, and edit point — then iterate one variable at a time. Writing prompts for video generation is closer to writing a shot list than writing a description.
More Posts

What Is MiniMax H3? What It Can Make, How Good It Is, What It Costs
MiniMax H3 (Hailuo 3) generates video and audio together in one pass. What kinds of clips it makes, what it can't do, how to start, and pricing from $0.08/second.


MiniMax H3 API Guide: 2K Video, Native Audio & Limits
MiniMax H3 API guide to 768P and 2K output, 4–15 second videos, native stereo audio, multimodal references, input limits, and async results.


MiniMax H3 Open Weights: Download & ComfyUI Guide
Download MiniMax H3 open weights from Hugging Face or ModelScope. 33B parameters, ComfyUI setup, and the license territory restrictions most guides skip.

Generate your first image with MiniMax H3 — right now
Reliable non-Latin text rendering, directed editing, and 50+ ready-to-use prompts. No downloads — just open in your browser.