MiniMax H3 Reference-to-Video Prompt Builder
When you hand MiniMax H3 reference images, clips or audio, it stops reading the three-field prompt the other modes use and expects six sections instead. This page writes those six, numbers every reference label for you, and catches the format mistakes that make a reference get quietly ignored.
Builder
Describe your reference assets, get the six sections
Add one row per file you plan to upload and say what each one is for. That single choice decides whether the file gets its own label line or is cited inside a subject definition — the distinction MiniMax H3 reference-to-video prompts get wrong most often. Start from a preset and edit from there.
Reference assets
2/6One row per uploaded file. Picture, Video and Audio numbers run independently, each in the order you add them.
Becomes a <Subject N> line, and the file label is cited inside it. It gets no line of its own — that is the rule people break most.
Becomes a <Subject N> line, and the file label is cited inside it. It gets no line of its own — that is the rule people break most.
Shots
2/3Playback order. Shot 1 carries no timestamp; later shots are stamped at their cut time.
Ref2VA prompt
344 words · band 350–500English, as H3 expects. Dialogue, lyrics and on-screen text keep their original language inside <d> tags.
subject_definitions: <Subject 1> is the young woman in <Picture 1>, with shoulder-length auburn hair, a grey knit sweater and a thin gold chain at the collar. <Subject 2> is the narrow bookshop interior in <Picture 2>, with floor-to-ceiling oak shelves, a brass desk lamp on the counter and a rolling ladder. summary: [reference generation] <Subject 1> takes a clothbound volume down from the upper shelves of <Subject 2> and carries it to the counter, where she settles and looks up toward the door. retention_analysis: <Subject 1> (appears in [Shot 1], [Shot 2]): fully_preserved - her face, auburn hair length, grey knit sweater and gold chain are carried into both shots without change. <Subject 2> (appears in [Shot 1], [Shot 2]): fully_preserved - the oak shelving, brass desk lamp and rolling ladder remain the only set, and no additional room or exit is introduced. detailed_description: The target video is in a warm, realistic style with soft practical lighting and a shallow depth of field. [Shot 1] A medium-wide shot establishes <Subject 2>, the narrow bookshop, with floor-to-ceiling oak shelves running the length of the frame and a brass desk lamp burning at the counter on the right. Late-afternoon light enters from a window off-frame left and falls in a soft band across the middle shelves, leaving the upper rows in shadow. <Subject 1>, the young woman with shoulder-length auburn hair and a grey knit sweater, stands on the second rung of the rolling ladder with her left hand flat against a shelf edge for balance. She tilts her head back to read the spines above her, then draws a clothbound volume free with two fingers, and the books on either side tip inward to fill the gap it leaves. She turns the volume over to read the spine, and a thin line of dust lifts off the top edge and drifts through the band of light. The camera pushes in with small amplitude at slow speed toward the book in her hands. [Shot 2] At 00:04.000, the shot cuts to a close-up of <Subject 1> at the counter of <Subject 2>, with the brass desk lamp just inside the right edge of frame throwing warm light across one side of her face and leaving the other in shadow. She sets the clothbound volume down on the wood, flattens the cover with her palm, and holds it there a moment longer than she needs to. Her thumb moves once across the corner of the boards. She lets out a short breath and her shoulders drop as it leaves her. The thin gold chain at her collar catches the lamp as she shifts her weight onto her other foot. She lifts her eyes from the book and looks up and past the camera toward the door, her expression settling into something closer to expectation than surprise. The camera holds a static shot throughout, letting her come to rest inside a frame that does not move with her. overall_soundscape: A quiet indoor room tone runs throughout, with the creak of the ladder under a shifting weight, pages riffling under a thumb, and the soft knock of board covers meeting the counter. non_diegetic_music: A sparse nylon-string guitar figure at a slow tempo, with a sustained low string underneath and no swell.
Opens the generator with this prompt loaded in reference-to-video mode. Add your files there.
Format check
No format problems found.
Summary prefix
Greyed-out types are set by your assets. Click the rest to add a relationship the assets do not imply.
The format
Six sections, not three
Text-to-video, image-to-video and the keyframe modes all share the same three core fields. As soon as a request carries reference assets, MiniMax H3 reference-to-video expects a different, longer structure — and the three-field prompt you used everywhere else is no longer the right shape.
| Section | What it carries |
|---|---|
subject_definitions | One line per tracked item, naming its label, what the label denotes and the features to follow. |
summary | One short paragraph, opened by a bracketed task-type prefix that states what kind of job this is. |
retention_analysis | One line per label, each carrying a fixed relationship marker and a note on what survives. |
detailed_description | The main body: shots in playback order, with labels inserted where they apply. |
overall_soundscape | Ambience and physical sound across the whole clip. Same meaning as in the base modes. |
non_diegetic_music | Score only the audience hears. Same meaning as in the base modes. |
The last two are unchanged from the base guide. The first four exist only in reference-to-video, which is why a prompt copied from a text-to-video template loses its references without any visible error.
Labels
Four label types, and the rule that decides which you get
Every reference-to-video asset is addressed by a label, and the same label keeps its meaning across all six sections. Picture, Video and Audio numbers are assigned independently, so <Video 1> and <Audio 2> can come from the same file.
<Subject N>Reusable visible content: a person, animal, object, scene, costume, prop, style, action or pose. One subject can draw on several files, and one file can supply several subjects.
<Picture N>An image acting as a concrete frame — first frame, keyframe, last frame — or as a storyboard anchor for shot planning.
<Video N>A whole-video relationship: the clip being edited, the clip being continued, or a source of cuts, pacing and camera structure.
<Audio N>An audio signal being copied or referenced, whether a standalone file or the synchronized track of a reference video.
The rule most reference-to-video prompts break
If an image exists only to define a character, scene, costume or style, it does not get a standalone <Picture N> entry. Its label is cited inside the <Subject N> definition instead, and it never appears in retention_analysis. A file earns its own line only when it plays a structural role — anchoring a frame, being edited, being continued, or carrying audio.
Wrong
<Picture 1> is a reference image of the young woman. <Subject 1> is the young woman.
Right
<Subject 1> is the young woman in <Picture 1>, with long dark hair, a blue cardigan and a thin silver necklace.
Retention markers
Two marker sets, and they are not interchangeable
Every line in retention_analysis carries a fixed English marker saying how strongly reference-to-video holds that asset. The visual track and the audio track use different sets — a distinction almost every third-party write-up drops, because the visual four are the ones people quote. Only weak_reference belongs to both.
For <Subject N>, <Picture N> and <Video N>
| Marker | Meaning | Typical use |
|---|---|---|
fully_preserved | The defined role of the reference is kept in full. | A character keeps face, hair and wardrobe across every shot. |
partially_preserved | Still used, but some defined characteristics change or are only partly retained. | A room keeps its layout while the wall treatment changes. |
attribute_transfer | Characteristics move onto a different identifiable target subject. | One person’s outfit is put on another person. |
weak_reference | Only broad similarity of style, category, composition or atmosphere survives. | A clip supplies pacing and nothing else. |
For <Audio N>
| Marker | Meaning | Typical use |
|---|---|---|
fully_copy | The complete source audio becomes the video’s complete final track. | An edit keeps the original interview audio untouched. |
partially_copy | Part of the timeline or selected layers are copied, or layers are added or replaced afterwards. | The dialogue is copied but the ambience is rebuilt. |
reference | Nothing is copied: only timbre, rhythm, music style, dialogue content or texture is referenced. | A speaker follows a reference voice without reusing it. |
weak_reference | Only broad similarity of category or atmosphere survives. | A track supplies a general mood. |
Pick a marker only inside the role that label already has in subject_definitions. New actions, backgrounds or events added in the target video are not losses of fidelity, so they do not justify downgrading a marker.
Summary prefix
Six task types, combined with a plus
Every reference-to-video summary opens with a bracketed prefix naming what kind of job this is. Choose by the role each asset actually plays, not by what kind of file it is — a reference video that only supplies pacing is reference generation, not video editing.
| Task type | When it applies |
|---|---|
keyframe completion | An image is a concrete frame anchor: first frame, keyframe, last frame or another fixed frame. |
reference generation | An asset guides generation — character, scene, style, action, camera, storyboard — without being a concrete frame or an edited source. |
video editing | An existing source video is directly modified. Generating between still keyframes does not count. |
video continuation | New content continues, extends, resumes or transitions from an existing source video. |
audio reuse | The same audio signal is reused in full or in part. |
audio reference | Only style, timbre, dialogue content, texture, beat or continuity is referenced, with no signal copied. |
Combining
Join types with a plus and never repeat one. Continuing a clip while using an image as the last frame is [video continuation + keyframe completion]. Editing a clip whose original audio stays audible is [video editing + audio reuse] — the audio half is easy to forget, because leaving it out breaks nothing visible.
One fixed sentence
For an edit, the summary must begin — right after the prefix — with “The target video is an edited version of <Video 1>.” The builder writes it for you as soon as an asset is marked as the source being edited.
Against the base modes
What reference-to-video changes, and what it leaves alone
Shots, camera vocabulary, speaker IDs and dialogue tags are shared with the base modes, so everything you already know about those still applies in reference-to-video. Four things move.
| Base modes | Reference-to-video | |
|---|---|---|
| Main field | integrated_multimodal_description | detailed_description |
| Style sentence | Written after [Shot 1] | Established on its own line before [Shot 1] |
| Reference labels | Not used | Inserted at first appearance and wherever the role applies |
| Audio | Describes the video’s own sound | Also states whether each cited signal is copied or referenced |
If your images are a first and last frame rather than character references, that is a different mode with a different alignment sentence — the first-and-last-frame builder writes that one instead.
Camera movement is written the same way in both, so the twenty motion types on the camera control page apply here unchanged.
Troubleshooting
When reference-to-video ignores your file
Reference-to-video fails quietly. The clip renders, the credits are spent, and the only symptom is that the model did not follow the file you uploaded. These are the format causes worth checking first.
A character drifts from the uploaded image
Likely cause
The image was given a standalone <Picture N> entry, so it reads as a frame anchor rather than a subject to reproduce.
Fix
Cite the image inside a <Subject N> definition and delete its separate line.
A reference is never applied at all
Likely cause
Its label is defined but never cited in detailed_description, so nothing tells the model where it takes effect.
Fix
Introduce each label at its first clear appearance and keep using it in later shots.
Reused audio comes out rebuilt rather than copied
Likely cause
The audio line carries a visual marker, or reference where the signal was meant to be copied.
Fix
Use fully_copy or partially_copy, and say so in the section matching the audible layer.
An edit loses its original dialogue
Likely cause
The prefix says video editing but not audio reuse, so nothing declares that the source audio stays.
Fix
Write [video editing + audio reuse] and add an <Audio N> line for the synchronized track.
The first frame is not reproduced
Likely cause
A single image was described as a reference rather than a frame anchor, which is image-to-video’s job, not a subject definition.
Fix
Mark it as the first frame of its shot so it becomes a <Picture N> line and adds keyframe completion.
None of this extends a clip past fifteen seconds — for that you plan segments and carry the last frame forward, which the long-video planner handles.
FAQ
Questions about reference-to-video prompts
Do I have to use all six sections?
Yes, once the request carries reference assets and you are in reference-to-video. The last two accept N/A when there is nothing to say, but the section still appears. If you have no reference files at all, you are not in this mode and should use the three-field prompt instead.
Should the prompt be in English?
Write the six reference-to-video sections in English. The one exception the guide makes is dialogue, lyrics and text visibly present in the scene, which keep their original language inside <d> tags — so a Chinese line stays Chinese and only the surrounding description is English.
How many reference images can I upload?
The reference-to-video mode of the generator on this site accepts up to three reference images. Numbering in the prompt follows the order you add them, so keep the builder rows in the same order as your uploads.
What is the difference between <Subject 1> and <Picture 1>?
<Picture 1> addresses the file. <Subject 1> addresses the content you want reused from it. If the file is only there to show what someone looks like, you want a subject, and the picture label is cited inside its definition rather than standing alone.
Can one subject come from two files?
Yes. The guide gives the case of appearance from an image and motion from a video, combined into one subject that names what each file supplies. Write that directly in the definition; the builder’s single-source rows cover the common case.
Why is there no (S1) in retention_analysis?
Speaker IDs are assigned by the order of vocal events in the target video and belong in the description. The guide states that they do not appear in retention_analysis at all, even when an audio reference is bound to a speaker — that binding is written in subject_definitions instead.
Does a reference video automatically create an <Audio N>?
No. A reference video containing sound does not create an audio label on its own. You add <Audio N> only when the track is actually copied or referenced, and Video and Audio numbers are independent, so the same file can be <Video 1> and <Audio 2>.
How long should detailed_description be?
The guide puts generation tasks at roughly 350 to 500 English words and asks that detail be spread across shots by information load. Dialogue-heavy work should fit the spoken timeline rather than hit a word count, and edits scale with the source clip instead.
Write the six reference-to-video sections, then run them
The builder hands the finished prompt to the generator in reference-to-video mode, where you attach the files the labels refer to.
Open the generatorFor the reasoning behind the format rather than the mechanics, the H3 prompt format guide walks through a full example end to end.
Ready-made prompts for the other modes live in the prompt library.
MiniMax H3 Full-Reference Mode Rewrite Output Format Guide · September 2026