MiniMax H3 First and Last Frame: How to Write FL2VA Prompts
2026/08/20

MiniMax H3 First and Last Frame: How to Write FL2VA Prompts

MiniMax H3's first and last frame mode needs an alignment line, one shot, and a written middle. The exact FL2VA format, six use cases, and why last frames miss.

The first time I ran a first-and-last-frame job on MiniMax H3, I wrote two careful paragraphs. One described the opening image in detail. One described the closing image in detail. What came back was eight seconds of crossfade — the two stills melting into each other with nothing in between that looked like it had weight.

The prompt wasn't badly written. It was answering a question the model hadn't asked. H3 can already see both images. Describing them back is the one thing an FL2VA prompt does not need to do.

FL2VA is one of four keyframe modes

MiniMax ships a prompt-writing guide with the open H3 weights, and it splits keyframe work into four modes that differ mainly in which images are pinned to which moment:

ModeImagesWhat's locked
T2VAnonenothing — the model invents framing and motion
I2VAfirst framethe opening
FL2VAfirst + last frameboth ends
L2VAlast framethe payoff

FL2VA is the one people reach for when the ending already exists: a product hero shot, an approved final design, a specific closing pose. You are not asking the model what happens. You are asking it to get from A to B convincingly.

That framing changes what belongs in the prompt. Everything the model can see is redundant. Everything between the two frames is yours to supply.

The alignment line most prompts skip

Before any of the content fields, an FL2VA prompt opens with a single instruction sentence that tells H3 which picture belongs to which moment:

How the reference pictures align with the target video — Picture 1 (from Shot 1) aligns with the 0.00-second mark of the target video; Picture 2 (from Shot N) aligns with the S.SS-second mark of the target video.

Three details in there are easy to get wrong:

N is the index of the actual final shot. Single-shot clip, and it is Shot 1 in both places. Add a cut and it becomes Shot 2. Leave it wrong and you have told the model your last frame belongs to a shot that never ends the video.

S.SS carries exactly two decimals. An eight-second clip is 8.00. Not 8, not 8.0. Nothing about the clip changes if you write it loosely, but the strict form is what the model was trained on.

One blank line follows. The instruction is its own paragraph, then the three content fields start.

I2VA and L2VA each use a different opening sentence — I2VA's is For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced. Copying an I2VA prompt and attaching a second image is the fastest way to have your end frame quietly ignored.

If you'd rather not hand-assemble any of this, the first and last frame prompt generator writes the alignment line from your duration and shot count, and flags the mistakes below while you type.

The body is a path, in four beats

After the alignment line come three fields, always in this order: integrated_multimodal_description, overall_soundscape, non_diegetic_music. The first one is where FL2VA is won or lost.

MiniMax's recommended structure for it is a four-beat path:

first-frame state → observable intermediate changes → progressively narrowing differences → last-frame state

Written out, a single shot looks like this:

[Shot 1] <style>, <what Picture 1 shows>, matching the position, framing, and
lighting established by Picture 1. The camera <motion> as <the intermediate
changes, in order>. Toward the end of the shot the differences narrow until
<what Picture 2 shows>, landing on the exact pose, spacing, and composition
established by Picture 2 at the <S.SS>-second mark.

The middle clause is the one that carries real information. "The hands lift the lid straight up and set it aside, then withdraw as the tissue paper falls open" gives the model three states to pass through. "The box opens to reveal the product" gives it one instruction and a lot of freedom, and freedom is what produced my crossfade.

Camera motion has its own three-part grammar: motion type, then amplitude, then speed, written as a natural action rather than stacked as tags. The camera pushes in with small amplitude at slow speed is right. Push In, small, slow is not. Medium amplitude and normal speed are the defaults and get left out.

One shot, unless you really mean it

The guide is direct about this: FL2VA favours a single shot so the model can interpolate continuously from the first frame to the last. Multiple shots should appear only when they are explicitly required.

This is the rule community prompts break most often, usually by importing habits from text-to-video, where cuts are free and often improve the result. In FL2VA a cut interrupts the only thing you asked the model to do.

If you do need one, the syntax is strict. The first shot takes no timestamp. Later shots open with a strictly increasing cut time inside the clip duration:

[Shot 2] At 00:03.500, the camera cuts to ...

And the last frame still has to be reached by the final shot at the end of the video, which is why the alignment line's shot index has to move with it.

For angle and distance changes, use camera motion instead. A cut should introduce genuinely new information — a new subject, space, viewpoint or time. Anything less is a job for a push-in.

Sound: two fields, one dividing line

overall_soundscape takes one to four sentences of ambience, physical action sound and non-verbal human sound across the whole clip. Rain, footsteps, fabric, impacts, breathing.

non_diegetic_music takes one to three sentences of score the characters cannot hear — instrumentation, tempo, rhythm, dynamics. N/A is a real answer and often the right one.

The dividing line is who can hear it. A radio playing in the room is diegetic and belongs in the body, alongside the action it happens over. Dialogue also lives in the body, wrapped in a tag with the speaker identified outside it:

The young woman with a quiet, breathy voice (S1) says: <d>[English] I get off at the next station.</d>

Dialogue length is a scheduling problem, not a writing one. Budget two to three words per second of clip and leave room for the action. A twelve-word line in a four-second clip has nowhere to go.

Six things people actually build with two frames

FL2VA earns its keep whenever the ending is already approved and the middle is the deliverable.

1. Product reveal. Closed packaging in frame one, product standing on its own in frame two. The whole job is keeping the label readable and the hands believable. Six seconds is usually enough; the temptation is to add a camera move that hides the product at the moment it appears.

2. Before and after. Renovation, makeover, restoration, a cleaned-up dataset. Both frames are the same subject from the same angle, which makes framing consistency easy and makes any drift in the middle very obvious. This is the mode's home turf.

3. Transformation and morph. One material or state into another — ice to water, sketch to sculpture, day to night on the same street. The failure mode here is the model choosing dissolve over transformation, which is again a symptom of an unwritten middle.

4. Logo and brand stings. Frame two is a fixed asset you cannot let the model redraw. Say so in the landing clause, and keep the clip short enough that there is no time for reinterpretation. If the logo still drifts, L2VA is worth trying instead: it pins only the closing frame, so the model is not also negotiating an opening image.

5. Character continuity across a cut. You have the last frame of shot A and the first frame of shot B from another tool, and you need the connective tissue. Here the two-shot syntax is legitimately the right call, because a cut is what you are trying to build.

6. Storyboard to animatic. Two consecutive boards become the motion between them. Because storyboards are drawn, style consistency comes almost free, and the useful output is timing information — you find out whether the beat fits in four seconds before anyone animates it.

Every one of these shares a shape: fixed opening, fixed payoff, and a transformation that has to be physically plausible. The pair does more work than the prompt does.

When the last frame refuses to land

Most failures trace back to one of a handful of causes.

Nothing converges. The prompt described two stills instead of the path. Write intermediate states, then say the differences narrow toward Picture 2.

Identity drifts halfway. Clothing, colour or a key prop was named only in the opening clause. Carry the anchors into the motion sentence so they stay fixed while the subject moves.

An unrequested cut appears. Multi-shot phrasing leaked into a single-shot job. Strip "cuts to" from the middle and use camera motion.

Motion rushes, then stalls. The action budget and the duration disagree. One clear transformation per clip; move the cut time or lengthen the clip rather than compressing the writing.

The pair is impossible. If no physical path connects the two frames — different lens, different room, different person — the model produces a dissolve because a dissolve is the only honest answer. Add an intermediate beat, lengthen the clip, or split it into two generations.

A three-step routine

Choose the pair before you write a word. Keep subject, lens and lighting consistent unless the change is the point, and check that a plausible path exists. A bad pair cannot be rescued by a good prompt.

Set the duration second. It appears verbatim in the alignment line and in the landing clause, so changing it later means editing the prompt in two places. Six to ten seconds covers most single-transformation clips.

Change one thing per retry. If the last frame misses, add intermediate states before you touch the camera. If motion smears, adjust the duration before you adjust the style. Running two changes at once tells you nothing about either.

Once the prompt is written, the same text works against the hosted API and a local ComfyUI graph, because the format comes from the guide shipped with the open weights. If you are running locally, LightX2V's multimodal prompt rewriter now has an 8B build on a Qwen3-VL base that reads first_frame and last_frame directly and emits the same three fields — around 9 GB of VRAM at Q4, which puts it within reach of the same card running H3. Pair it with the Turbo step-count guide if you are sampling in 4 to 8 steps rather than 20.

The bottom line

An FL2VA prompt is not a description of two pictures. It is a description of the trip between them, plus one sentence of bookkeeping that tells the model where each picture sits on the timeline.

Get the alignment line right, keep it to one shot, and spend your writing on the middle. The crossfade I got on my first attempt wasn't the model failing to understand the images — it was the model doing exactly what a prompt with no middle asks for.

Build the prompt in the FL2VA prompt generator, then run it with both frames attached in the H3 generator. For the full six-part format across reference-to-video and the other modes, the MiniMax H3 prompt guide covers the rest.

Free to try

Generate your first video with MiniMax H3 — right now

Create videos from text or a reference image, with ready-to-use prompt examples to help you get started. No downloads — just open it in your browser.