MiniMax H3 提示词指南:六类创作场景写法详解
2026/08/05

MiniMax H3 提示词指南:六类创作场景写法详解

MiniMax H3 视频生成提示词怎么写?三种输入模式选型、可复用的提示词公式、六类创作场景完整写法——看完就能上手。

大部分 H3 生成失败不是模型能力不行,是提示词没写对。模型老老实实执行了你给的指令——问题出在指令本身。

8 月 3 日 MiniMax 在发布开源权重的同时放出了官方提示词指南。这篇文章把那份指南提炼成一套选择框架、一个可复用的提示词模板、六类创作场景的完整写法。如果还没开通 API,先看 MiniMax H3 API 指南

选对输入模式

H3 支持三种输入方式,每种对提示词的要求不一样。选错模式是白白浪费生成次数的常见原因。

文生视频

不上传任何文件,全靠提示词描述。

适用场景:从剧本或脑中构思出发,手边没有可用的参考图片或视频。因为模型什么都看不到,每一个视觉细节都要在提示词里写清楚——人物外观、场景环境、动作、运镜、光线、声音设计。

文生视频必须手动指定画面比例(21:9、16:9、4:3、1:1、3:4 或 9:16),模型无法从不存在的图片推断。

首尾帧(图生视频)

上传起始帧、结束帧,或两者都传。提示词负责描述两帧之间发生了什么。

适用场景:产品展示(产品样子不能走形)、角色状态变化(站→坐、白天→夜晚)、场景转换、把静态海报动起来。图片锁住关键画面,提示词填运动和声音。

输出比例由输入图片决定,不用另外设置。首尾帧和 reference 类型素材不能混在同一个请求里。

全模态参考(参考生视频)

最灵活的模式。上传图片(最多 9 张)、视频(最多 3 段,总时长不超过 15 秒)、音频(最多 3 段,总时长不超过 15 秒),然后在提示词里给每份素材分配角色。

适用场景:手上有明确的资产需要保留——人物照片的脸、某段视频的运镜轨迹、某段录音的嗓音或环境氛围。提示词变成一份"选角表":告诉模型每份素材贡献什么。

注意:音频不能作为唯一参考,至少要搭配一张图片或一段视频。

怎么选

手边没有任何视觉素材,纯粹从文字出发 → 文生视频

有一两张关键图片必须原样出现在成片里(产品图、人物肖像、背景底板)→ 首尾帧

有多份参考资产(人物照片、运镜参考、配音录音、风格图)需要模型综合使用 → 全模态参考

一条经验:参考素材越多,提示词里的视觉描述可以越少——但"角色分配"文字要写得越多。文生视频的提示词长在描写,全模态参考的提示词长在调度。

提示词公式

MiniMax 官方指南把一条完整提示词拆成三部分:

完整提示词 = 参考素材说明 + 核心创意 + 视觉过程描写

三种输入模式通用,侧重不同。

第一部分:参考素材说明

上传了文件就要编号、标注角色。模型需要知道每份文件干什么用。

写法示意:@image 1 是角色参考——保留人脸、发型和服装。@video 1 提供运镜——跟着它推镜头的节奏。@audio 1 是环境音轨——匹配它的节奏感。

要点在于指出需要保留的具体特征。"用这张图做参考"太笼统。"用这张图做角色的脸和穿搭参考,但不要用它的背景"才是模型可以执行的约束。

文生视频不需要这段。首尾帧模式下,这里写清楚哪张图是起始帧、哪张是结束帧。

第二部分:核心创意

一句话概括主体、场景、事件、风格和运镜方向。相当于这条视频的主旨句。

示例:一位穿深蓝色大衣的女性在黄昏时分穿过淋湿的巷子,中景跟拍从背后缓慢推进,前方门洞透出暖色钨丝灯光。

如果这一段需要超过两句话才能说清,你大概率在描述多个镜头。拆开。

第三部分:视觉过程描写

按时间轴或分镜逐段写清楚每段画面里发生什么。

示例:

  • 0–3 秒:双手特写,调整衣领。机位不动。
  • 3–6 秒:拉开到中景。角色转身朝灯光方向走去。镜头缓慢跟进。
  • 6–10 秒:切到前方门洞。暖光洒在湿地面上。角色从画面左侧入画。

大多数提示词在这一段拉胯。写了一个氛围语("充满电影感和情感")就当描述完了。模型执行的是动作,不是形容词。

声音控制也放在这里。可以指定背景音乐风格、环境声、台词对白,或者用 "non-narrative music: N/A" 直接关掉背景音乐。

可复用模板

[参考素材说明]
@image 1:[角色——需要保留的特征]
@image 2:[角色——需要保留的特征]
@video 1:[角色——需要提取的元素]
@audio 1:[角色——需要匹配的内容]

[核心创意]
[主体] 在 [场景] 中 [动作],以 [运镜风格] 拍摄,[光线/氛围]。

[视觉过程]
0–[X]s:[景别]。[动作]。[运镜]。[声音]。
[X]–[Y]s:[景别]。[动作]。[运镜]。[声音]。
[Y]–[Z]s:[景别]。[动作]。[运镜]。[声音]。

[声音设计]
BGM:[风格/情绪,或 "N/A"]
环境声:[具体声音]
台词:[内容,如有]

不适用的部分直接删掉。文生视频把参考说明删掉,视觉描写写得更细。首尾帧把参考说明换成帧分配。

六类创作场景

以下六类场景来自 MiniMax 官方指南,这里扩展了可以直接改写套用的提示词写法。

1. 品牌片与电影感表达

目标是镜头语言——有动机的剪辑、有设计的转场、多镜头之间的视觉连贯。

这类通常用首尾帧模式。关键帧锁住每个镜头的开头和结尾,提示词负责填写镜头内部的运动和镜头之间的转场。硬切、黑屏、闪白、有动机的推镜头,全在视觉过程里写。

H3 提示词常用的运镜词汇:push-in(推镜)、pull-back(拉镜)、slow pan left(缓慢左摇)、tracking shot(跟踪镜头)、rack focus(焦点转移)、Dutch angle(荷兰角)、bird's eye(鸟瞰)、worm's eye(仰拍)、Steadicam follow(稳定器跟拍)、whip pan(甩镜)、static locked-off(固定机位)。

示例提示词(文生视频):

A man in a dark wool overcoat stands at the edge of a rooftop at
golden hour. City skyline behind him, backlit.

0–3s: Extreme close-up of his eyes. Shallow depth of field. The
city lights reflect in his pupils. Camera holds still. Sound: wind,
distant traffic.
3–5s: Black screen, 0.5s. Then cut to wide shot from behind —
the man silhouetted against the skyline. Camera slowly pushes in.
5–8s: Medium shot from the side. He turns his head to the right.
Key light from the setting sun, warm amber. Camera tracks around
to face him.
8–10s: Close-up of his hand releasing a paper airplane. Camera
follows it off the rooftop into the sky. Sound: paper flutter, then
silence.

BGM: minimal piano, single repeating note
Ambient: rooftop wind, faint city hum

2. 视觉创意与内容包装

走风格化路线——复古动画、硬边剪影、漫画拼贴、不对称分屏。提示词靠具体的视觉语言驱动风格,而不是靠风格标签。

这类适合用全模态参考模式。上传一张风格参考图和一段运镜参考视频,提示词分配角色。

示例提示词(全模态参考):

@image 1: style reference — preserve the flat color palette, thick
ink outlines, and halftone dot texture.
@video 1: motion reference — match the timing of its zoom-outs
and the rhythmic cuts on every beat.

A street musician plays saxophone under a bridge at night, drawn
in retro anime style with visible ink grain.

0–3s: Asymmetric split screen — left panel shows hands on keys
in close-up, right panel shows the full figure from a low angle.
3–5s: Panels merge into single frame. Camera pulls back to
reveal the bridge and the river below. Hard-edge silhouette
transition.
5–8s: Comic-panel collage — four rapid cuts of saxophone
close-ups, each held for 0.5s, framed with thick black borders.
8–10s: Return to single wide shot. Text "MIDNIGHT SET" appears
bottom-center in hand-drawn lettering. Hold 2s.

BGM: jazz saxophone, synced to the visual rhythm
Ambient: river water, echo under the bridge

3. 品牌与时尚内容

产品通过角色的动作自然进入画面,而不是静态展示。镜头跟着人走,产品是因为角色使用它、拿起它、穿着它才被看到的。

产品图作为 image reference 上传,角色描述要精确。提示词重点写产品如何自然地出现在场景动作中。

示例提示词(全模态参考):

@image 1: character reference — preserve face, hairstyle, and
body proportions.
@image 2: product reference — this is a leather crossbody bag in
cognac brown. Preserve the color, buckle detail, and strap width.

A woman walks along a desert highway at sunset, wearing a linen
shirt and wide-leg trousers. She carries the bag from @image 2 on
her left shoulder.

0–3s: Wide shot from the front. She walks toward camera. A
vintage car is parked on the roadside behind her. Golden hour
light, long shadows.
3–6s: Medium shot from the side. Her hand adjusts the bag strap.
Camera pans to follow. The bag fills the lower third of the frame.
6–9s: Close-up of the bag — the buckle catches sunlight. Shallow
depth of field. Camera holds, slight tilt up to her face. She
glances down and smiles.
9–10s: Pull back to wide. She continues walking. Dust rises from
her steps.

BGM: acoustic guitar, unhurried
Ambient: desert wind, gravel footsteps
Dialogue: none

4. 动画与风格化角色

角色参考图锁定人脸、发型、身材和服装。模型在整条片子里保持这些特征不漂。适合角色 PV、游戏 CG、动画 PV、IP 衍生内容。

提示词的关键技巧是显式枚举特征:把模型必须保持的每项属性列出来,而不是写"保持角色一致"。

示例提示词(全模态参考):

@image 1: character reference — preserve exactly: silver-white
bob-cut hair, violet eyes, sharp jawline, black turtleneck, silver
pendant necklace.

A character with @image 1's appearance stands in a neon-lit
cyberpunk alley. Rain falls. Reflections on wet ground.

0–4s: Medium shot, frontal. She looks directly at camera. Rain
hits her shoulders. Neon signs reflect on her face — alternating
pink and blue. Camera slowly pushes in.
4–7s: Cut to profile view. She raises her right hand. A
holographic interface appears from her palm. Camera tracks from
her face to the hologram.
7–10s: Over-the-shoulder shot looking at the hologram. Data
scrolls across it. Camera slowly rotates around to face her
through the translucent display.

BGM: ambient synth, low and pulsing
Ambient: rain, neon buzz, distant sirens

5. 产品与电商营销

从产品图出发,用提示词把静态照片变成 360 度展示、材质特写、人体工学演示、真实场景植入。

根据对运镜控制的需要选模式:单张产品图做首帧适合简单的旋转展示,多角度产品图做 reference 适合复杂的产品揭幕。

示例提示词(首帧模式):

First frame: product photo — a wireless ergonomic mouse on a
white surface, viewed from a 45-degree angle.

The mouse sits on a clean white desk. The camera showcases the
product from multiple angles.

0–3s: Camera slowly orbits the mouse clockwise, 90 degrees.
Clean white background. Soft diffused light from above.
3–5s: Close-up of the side grip texture. Camera pushes in until
the surface pattern fills the frame. Sound: soft click of the side
button.
5–8s: Pull back. A hand enters from the right and naturally grips
the mouse. Medium shot showing the ergonomic curve against the
palm. The desk surface changes to dark walnut.
8–10s: Wide shot of a full desk setup — monitor, keyboard, and
the mouse in use. The hand clicks and scrolls. Warm office
lighting, shallow depth of field on the mouse.

BGM: N/A
Ambient: quiet office, keyboard tapping, mouse click

6. 角色、物体与场景编辑

绿幕抠除、背景替换、台词替换。提示词要说清什么保留、什么替换。

背景替换的核心是光线匹配——如果角色身上主光从左边打来,新背景的光源也要在左边。光线方向不一致是最常见的穿帮。

台词替换的做法是把目标音频作为 reference 上传,提示词要求模型调整角色的口型和表演来匹配新音频。

示例提示词(全模态参考,背景替换):

@image 1: character on green screen — preserve the character's
appearance, pose, and left-side key lighting exactly.
@image 2: new background — a Japanese garden in autumn, warm
afternoon light from the left.

Place the character from @image 1 into the scene from @image 2.
Match the lighting direction — key light from the left in both.
Remove the green screen entirely.

0–5s: Medium shot. Character stands in the garden. Maple leaves
drift down around her. Gentle camera sway, handheld feel. Sound:
birdsong, light wind through leaves.
5–10s: Slow push-in to close-up. Leaves continue falling. One
leaf lands on her shoulder. She brushes it off. The lighting on
her face matches the garden's warm afternoon tones.

BGM: traditional koto, very quiet
Ambient: garden birds, wind, rustling leaves
Dialogue: none

常见错误

用氛围词代替动作。"充满电影感"不是指令。"镜头在 3 秒内缓慢推进,角色垂下目光"才是。

上传了素材但不标注角色。 传了三张图片,没说哪张是角色、哪张是风格、哪张是背景。模型随机取用。

参考素材互相打架。 两张不同人的脸。模型取平均,出来的人两边都不像。

写了一堆画质标签。 "4K""hyperrealistic""8K ultra HD"这些不映射到模型参数上,纯粹浪费提示词空间。

一次生成塞太多东西。 10 秒片子里塞五次场景切换、三个角色、一段对白。简化到每次生成一个清晰的戏剧节拍,生成完再剪到一起。

不管结束帧。 如果这条片子要跟其他片子剪在一起,写清楚最后一帧停在什么画面上。结尾没规划好,后期剪辑会很痛苦。

延伸阅读

上面的模板覆盖了大部分制作场景。从一个镜头开始,用四项检查验证成片——角色一致性、动作完整度、音画同步、剪辑点——然后每次只改一个变量重新生成。写 AI 视频提示词更像写分镜表,不像写作文。

免费试用

立即用 MiniMax H3 生成你的第一张图像

可靠的非拉丁文本渲染、精准的定向编辑,以及 50+ 开箱即用的提示词。无需下载——在浏览器中打开即可使用。