MiniMax H3 · FL2VA prompt builder

MiniMax H3 first and last frameprompt generator

H3 can already see both of your frames. What it cannot see is the path between them — and that is the only thing an FL2VA prompt is for. Describe the change, and this builder writes it in the exact format MiniMax ships with the model: alignment line, shot timeline, soundscape, score.

Official FL2VA syntaxEight live prompt checksRuns on hosted H3 or ComfyUI

Build the prompt

Two frames in, one motion path out

Start from a preset or write your own. Everything on the right is assembled live, including the alignment line most hand-written FL2VA prompts get wrong.

FL2VA prompt

Official MiniMax format — alignment line, then three fields.

How the reference pictures align with the target video — Picture 1 (from Shot 1) aligns with the 0.00-second mark of the target video; Picture 2 (from Shot 1) aligns with the 6.00-second mark of the target video.

integrated_multimodal_description: [Shot 1] Live-action, cinematic, a matte black gift box sits closed and centred on a pale marble counter under soft window light, its lid seam facing camera, matching the position, framing, and lighting established by Picture 1. The camera pushes in with small amplitude at slow speed as two hands enter from the lower edge, lift the lid straight up and set it aside, then withdraw as the tissue paper falls open around a glass serum bottle. Toward the end of the shot the differences narrow until the serum bottle stands upright and unobstructed in the centre of the counter with the opened box behind it and the printed label facing camera, landing on the exact pose, spacing, and composition established by Picture 2 at the 6.00-second mark.

overall_soundscape: A quiet room tone carries throughout, broken by the soft friction of the lid sliding free, the rustle of tissue paper, and the light tap of glass settling on stone.

non_diegetic_music: Sparse marimba notes at a slow tempo, joined by a warm synth pad that swells once as the bottle is revealed.
Run it on hosted MiniMax H3

Opens the image-to-video generator with this prompt loaded. Attach your start frame and end frame there.

Prompt check

Nothing flagged. Both frames are described, the middle has a path, and the clip is a single continuous shot.

Anatomy

What an FL2VA prompt is made of

Four parts, in a fixed order. Skip the alignment line and H3 has no idea which image is the opening and which is the payoff.

Alignment line

Always first

One sentence that pins Picture 1 to 0.00 seconds and Picture 2 to the end of the clip. The end time carries exactly two decimals, and a blank line follows before the fields start.

integrated_multimodal_description

Required

The body. Style, opening composition, the motion path between the frames, camera work, speakers, dialogue and any sound that happens on screen. This is where FL2VA is won or lost.

overall_soundscape

Required

One to four sentences of ambient, physical and non-verbal human sound across the whole clip. No dialogue, no singing, no music the characters can hear — those belong in the body.

non_diegetic_music

Required

One to three sentences on score the characters cannot hear. Instrumentation, tempo, rhythm, dynamics. N/A when there is none, which is a real answer rather than a cop-out.

MiniMax's own FL2VA example

An eight-second single shot from the prompt-writing guide shipped with the H3 weights, trimmed here for length. Note that the body never describes the two pictures — it describes the umbrella opening between them.

How the reference pictures align with the target video — Picture 1 (from Shot 1) aligns with the 0.00-second mark of the target video; Picture 2 (from Shot 1) aligns with the 8.00-second mark of the target video.

integrated_multimodal_description: [Shot 1] Live-action, cinematic, a rain-soaked cyclist begins in the position and framing established by Picture 1 …

overall_soundscape: Rain falls steadily on the pavement, followed by the metallic click of the umbrella runner …

non_diegetic_music: N/A
Read the full guide on Hugging Face

Where it goes wrong

Six ways the last frame misses

Nearly every failed first-and-last-frame clip is one of these. Five of the six are prompt problems, not model problems.

The last frame never arrives

Why: The prompt described two still images instead of the path between them, so the model had no reason to converge on anything.

Fix: Write the middle. Name the intermediate states the subject passes through, then say the differences narrow until the composition matches Picture 2 at the end of the final shot.

Identity or wardrobe drifts halfway

Why: Nothing in the body told the model which attributes are fixed, so it re-invented them once the subject started moving.

Fix: Carry the anchors — clothing, colour, key props, spatial relationships — into the motion sentence rather than leaving them in the opening clause only.

A cut appears you never asked for

Why: Multi-shot phrasing leaked into a single-shot job, or the description implied a viewpoint jump the model resolved as an edit.

Fix: Keep FL2VA to one shot unless you genuinely need a cut. Use camera motion for distance and angle changes; save "the camera cuts to" for new information.

Motion rushes then stalls

Why: The clip length and the amount of described action disagree. Six seconds of writing crammed into a four-second render has to go somewhere.

Fix: Match the action budget to the duration, or move the cut time. One clear transformation per clip beats three that all get clipped.

Dialogue runs past the end of the clip

Why: The line is longer than the duration allows at natural speaking pace, so it is truncated or the mouth keeps moving after the audio stops.

Fix: Budget roughly two to three words per second of screen time, and leave room for the action. Long lines want a longer clip, not faster delivery.

Music and ambience fight each other

Why: Diegetic sound was written into non_diegetic_music, or the soundscape repeated what the body already described.

Fix: Sound the characters can hear goes in the body. Score only the audience hears goes in non_diegetic_music. The soundscape summarises ambience, impacts and breath — nothing else.

Pick the mode

FL2VA against the other three

H3 takes keyframes four different ways. The alignment line is what changes between them, and getting it wrong is the fastest way to have your reference images quietly ignored.

ModeImagesAlignment lineControlPick it when
T2VAText onlyNo alignment lineLowestYou have no reference art and want the model to invent the framing as well as the motion.
I2VAFirst framePicture 1 at 0.00sStart lockedThe opening image matters and you are happy to let the model decide where the shot ends up.
FL2VAFirst + last framePicture 1 at 0.00s, Picture 2 at the endBoth ends lockedYou already know both the opening and the closing image and need a believable path between them.
L2VALast framePicture 1 at the endEnd lockedThe payoff image is fixed — a logo, a reveal, a final pose — and the run-up is yours to invent.

Use cases

What people actually build with two frames

Each preset in the builder is one of these. They share a shape: a fixed opening image, a fixed payoff image, and a transformation that has to be believable in between.

E-commerce

Product reveal

Closed packaging in frame one, the product standing on its own in frame two. The hardest part is keeping the label readable the whole way through.

6s · single shot · Live-action, cinematic

Renovation / makeover

Before and after

The classic FL2V job. The two frames are the same room or the same face, and the model has to invent a believable path between them.

8s · single shot · Live-action

Creative process

Sketch to final

Frame one is the rough, frame two is the finished piece. Works for illustration, design, lettering and 3D turntables.

10s · single shot · 2D-animated

How to use it

From two stills to a finished clip

  1. 1

    Pick the two frames first

    FL2VA is only as good as the pair you feed it. Keep the subject, lens and lighting consistent between the two images unless the change itself is the point, and make sure a plausible physical path exists between them.

  2. 2

    Set the duration before you write

    The end time appears verbatim in the alignment line, so changing it later means rewriting the prompt. Four to fifteen seconds is the supported range; six to ten covers most single-transformation clips.

  3. 3

    Describe the path, not the pictures

    The model can already see both frames. Spend the body on what moves, what changes hands, how the composition evolves, and how the differences narrow toward the last frame.

  4. 4

    Copy the prompt and run both frames

    Paste the generated prompt into the H3 image-to-video generator with your start image and end image attached, or into a local ComfyUI graph. Judge the result on whether the final frame actually lands, not on how the middle looks.

    Open the H3 generator

FAQ

First and last frame questions

What is a MiniMax H3 first and last frame prompt?

It is the FL2VA prompt format. The prompt opens with an alignment line that pins Picture 1 to the 0.00-second mark and Picture 2 to the end of the clip, then supplies three fields: integrated_multimodal_description, overall_soundscape and non_diegetic_music. H3 already sees both images, so the body describes the motion path between them rather than the images themselves.

Why does the alignment line need two decimal places?

MiniMax specifies the end time as S.SS — exactly two decimals. An eight-second clip is written 8.00, not 8 or 8.0. It is a formatting rule rather than a precision one, and this builder writes it for you.

Should an FL2VA prompt use one shot or several?

One, in almost every case. MiniMax states that FL2VA favours a single shot so the model can interpolate continuously from the first frame to the last, and that multiple shots should only appear when they are explicitly required. If you do add a cut, the last frame must be reached by the final shot at the end of the video.

Can I generate a first-and-last-frame video on this site?

Yes. The image-to-video generator accepts a start image and an optional end image, and passes them to H3 as first_frame_url and last_frame_url. Build the prompt here, then open the generator with both frames attached.

How is FL2VA different from image-to-video?

Image-to-video (I2VA) locks only the opening frame and lets the model decide where the shot ends. FL2VA locks both ends, which is what you want when the closing image is already fixed — a product hero shot, a finished design, a specific final pose. The trade-off is that a physically implausible pair produces a worse clip than a single frame would.

Does the prompt this tool writes work in ComfyUI too?

It does. The format comes from the prompt-writing guide shipped with the open H3 weights, so the same text works against the hosted API and against a local ComfyUI graph. The LightX2V multimodal rewriter, now distilled to an 8B Qwen3-VL base that reads first_frame and last_frame directly, emits the same three fields.

How long should the dialogue be?

Budget two to three words per second of clip and leave room for the action. Dialogue goes inside a <d>[Language] ...</d> tag with the speaker identified outside it, and this builder warns you when a line looks too long for the duration you set.

What if the two frames are too far apart?

The model will invent something, and usually it will look like a dissolve rather than motion. Either add an intermediate beat to the motion path so the change has visible stages, lengthen the clip, or split the idea across two generations and cut them together.

You have both frames. Write the middle.

The generator on this site takes a start frame and an end frame and hands both to MiniMax H3. Build the prompt above, open the generator, attach the pair, and see whether your last frame actually lands.

The FL2VA syntax on this page — the alignment sentence, the [Shot N] cut format, the camera vocabulary and the three output fields — is taken from MiniMaxAI's Video Prompt Writing Guide published with the open H3 weights, verified August 2026. This site is not affiliated with MiniMax.