
MiniMax H3 提示词指南:六类创作场景写法详解
MiniMax H3 视频生成提示词怎么写?三种输入模式选型、可复用的提示词公式、六类创作场景完整写法——看完就能上手。
大部分 H3 生成失败不是模型能力不行,是提示词没写对。模型老老实实执行了你给的指令——问题出在指令本身。
8 月 3 日 MiniMax 在发布开源权重的同时放出了官方提示词指南。这篇文章把那份指南提炼成一套选择框架、一个可复用的提示词模板、六类创作场景的完整写法。如果还没开通 API,先看 MiniMax H3 API 指南。
选对输入模式
H3 支持三种输入方式,每种对提示词的要求不一样。选错模式是白白浪费生成次数的常见原因。
文生视频
不上传任何文件,全靠提示词描述。
适用场景:从剧本或脑中构思出发,手边没有可用的参考图片或视频。因为模型什么都看不到,每一个视觉细节都要在提示词里写清楚——人物外观、场景环境、动作、运镜、光线、声音设计。
文生视频必须手动指定画面比例(21:9、16:9、4:3、1:1、3:4 或 9:16),模型无法从不存在的图片推断。
首尾帧(图生视频)
上传起始帧、结束帧,或两者都传。提示词负责描述两帧之间发生了什么。
适用场景:产品展示(产品样子不能走形)、角色状态变化(站→坐、白天→夜晚)、场景转换、把静态海报动起来。图片锁住关键画面,提示词填运动和声音。
输出比例由输入图片决定,不用另外设置。首尾帧和 reference 类型素材不能混在同一个请求里。
全模态参考(参考生视频)
最灵活的模式。上传图片(最多 9 张)、视频(最多 3 段,总时长不超过 15 秒)、音频(最多 3 段,总时长不超过 15 秒),然后在提示词里给每份素材分配角色。
适用场景:手上有明确的资产需要保留——人物照片的脸、某段视频的运镜轨迹、某段录音的嗓音或环境氛围。提示词变成一份"选角表":告诉模型每份素材贡献什么。
注意:音频不能作为唯一参考,至少要搭配一张图片或一段视频。
怎么选
手边没有任何视觉素材,纯粹从文字出发 → 文生视频。
有一两张关键图片必须原样出现在成片里(产品图、人物肖像、背景底板)→ 首尾帧。
有多份参考资产(人物照片、运镜参考、配音录音、风格图)需要模型综合使用 → 全模态参考。
一条经验:参考素材越多,提示词里的视觉描述可以越少——但"角色分配"文字要写得越多。文生视频的提示词长在描写,全模态参考的提示词长在调度。
提示词公式
MiniMax 官方指南把一条完整提示词拆成三部分:
完整提示词 = 参考素材说明 + 核心创意 + 视觉过程描写
三种输入模式通用,侧重不同。
第一部分:参考素材说明
上传了文件就要编号、标注角色。模型需要知道每份文件干什么用。
写法示意:@image 1 是角色参考——保留人脸、发型和服装。@video 1 提供运镜——跟着它推镜头的节奏。@audio 1 是环境音轨——匹配它的节奏感。
要点在于指出需要保留的具体特征。"用这张图做参考"太笼统。"用这张图做角色的脸和穿搭参考,但不要用它的背景"才是模型可以执行的约束。
文生视频不需要这段。首尾帧模式下,这里写清楚哪张图是起始帧、哪张是结束帧。
第二部分:核心创意
一句话概括主体、场景、事件、风格和运镜方向。相当于这条视频的主旨句。
示例:一位穿深蓝色大衣的女性在黄昏时分穿过淋湿的巷子,中景跟拍从背后缓慢推进,前方门洞透出暖色钨丝灯光。
如果这一段需要超过两句话才能说清,你大概率在描述多个镜头。拆开。
第三部分:视觉过程描写
按时间轴或分镜逐段写清楚每段画面里发生什么。
示例:
- 0–3 秒:双手特写,调整衣领。机位不动。
- 3–6 秒:拉开到中景。角色转身朝灯光方向走去。镜头缓慢跟进。
- 6–10 秒:切到前方门洞。暖光洒在湿地面上。角色从画面左侧入画。
大多数提示词在这一段拉胯。写了一个氛围语("充满电影感和情感")就当描述完了。模型执行的是动作,不是形容词。
声音控制也放在这里。可以指定背景音乐风格、环境声、台词对白,或者用 "non-narrative music: N/A" 直接关掉背景音乐。
可复用模板
[参考素材说明]
@image 1:[角色——需要保留的特征]
@image 2:[角色——需要保留的特征]
@video 1:[角色——需要提取的元素]
@audio 1:[角色——需要匹配的内容]
[核心创意]
[主体] 在 [场景] 中 [动作],以 [运镜风格] 拍摄,[光线/氛围]。
[视觉过程]
0–[X]s:[景别]。[动作]。[运镜]。[声音]。
[X]–[Y]s:[景别]。[动作]。[运镜]。[声音]。
[Y]–[Z]s:[景别]。[动作]。[运镜]。[声音]。
[声音设计]
BGM:[风格/情绪,或 "N/A"]
环境声:[具体声音]
台词:[内容,如有]不适用的部分直接删掉。文生视频把参考说明删掉,视觉描写写得更细。首尾帧把参考说明换成帧分配。
六类创作场景
以下六类场景来自 MiniMax 官方指南,这里扩展了可以直接改写套用的提示词写法。
1. 品牌片与电影感表达
目标是镜头语言——有动机的剪辑、有设计的转场、多镜头之间的视觉连贯。
这类通常用首尾帧模式。关键帧锁住每个镜头的开头和结尾,提示词负责填写镜头内部的运动和镜头之间的转场。硬切、黑屏、闪白、有动机的推镜头,全在视觉过程里写。
H3 提示词常用的运镜词汇:push-in(推镜)、pull-back(拉镜)、slow pan left(缓慢左摇)、tracking shot(跟踪镜头)、rack focus(焦点转移)、Dutch angle(荷兰角)、bird's eye(鸟瞰)、worm's eye(仰拍)、Steadicam follow(稳定器跟拍)、whip pan(甩镜)、static locked-off(固定机位)。
示例提示词(文生视频):
A man in a dark wool overcoat stands at the edge of a rooftop at
golden hour. City skyline behind him, backlit.
0–3s: Extreme close-up of his eyes. Shallow depth of field. The
city lights reflect in his pupils. Camera holds still. Sound: wind,
distant traffic.
3–5s: Black screen, 0.5s. Then cut to wide shot from behind —
the man silhouetted against the skyline. Camera slowly pushes in.
5–8s: Medium shot from the side. He turns his head to the right.
Key light from the setting sun, warm amber. Camera tracks around
to face him.
8–10s: Close-up of his hand releasing a paper airplane. Camera
follows it off the rooftop into the sky. Sound: paper flutter, then
silence.
BGM: minimal piano, single repeating note
Ambient: rooftop wind, faint city hum2. 视觉创意与内容包装
走风格化路线——复古动画、硬边剪影、漫画拼贴、不对称分屏。提示词靠具体的视觉语言驱动风格,而不是靠风格标签。
这类适合用全模态参考模式。上传一张风格参考图和一段运镜参考视频,提示词分配角色。
示例提示词(全模态参考):
@image 1: style reference — preserve the flat color palette, thick
ink outlines, and halftone dot texture.
@video 1: motion reference — match the timing of its zoom-outs
and the rhythmic cuts on every beat.
A street musician plays saxophone under a bridge at night, drawn
in retro anime style with visible ink grain.
0–3s: Asymmetric split screen — left panel shows hands on keys
in close-up, right panel shows the full figure from a low angle.
3–5s: Panels merge into single frame. Camera pulls back to
reveal the bridge and the river below. Hard-edge silhouette
transition.
5–8s: Comic-panel collage — four rapid cuts of saxophone
close-ups, each held for 0.5s, framed with thick black borders.
8–10s: Return to single wide shot. Text "MIDNIGHT SET" appears
bottom-center in hand-drawn lettering. Hold 2s.
BGM: jazz saxophone, synced to the visual rhythm
Ambient: river water, echo under the bridge3. 品牌与时尚内容
产品通过角色的动作自然进入画面,而不是静态展示。镜头跟着人走,产品是因为角色使用它、拿起它、穿着它才被看到的。
产品图作为 image reference 上传,角色描述要精确。提示词重点写产品如何自然地出现在场景动作中。
示例提示词(全模态参考):
@image 1: character reference — preserve face, hairstyle, and
body proportions.
@image 2: product reference — this is a leather crossbody bag in
cognac brown. Preserve the color, buckle detail, and strap width.
A woman walks along a desert highway at sunset, wearing a linen
shirt and wide-leg trousers. She carries the bag from @image 2 on
her left shoulder.
0–3s: Wide shot from the front. She walks toward camera. A
vintage car is parked on the roadside behind her. Golden hour
light, long shadows.
3–6s: Medium shot from the side. Her hand adjusts the bag strap.
Camera pans to follow. The bag fills the lower third of the frame.
6–9s: Close-up of the bag — the buckle catches sunlight. Shallow
depth of field. Camera holds, slight tilt up to her face. She
glances down and smiles.
9–10s: Pull back to wide. She continues walking. Dust rises from
her steps.
BGM: acoustic guitar, unhurried
Ambient: desert wind, gravel footsteps
Dialogue: none4. 动画与风格化角色
角色参考图锁定人脸、发型、身材和服装。模型在整条片子里保持这些特征不漂。适合角色 PV、游戏 CG、动画 PV、IP 衍生内容。
提示词的关键技巧是显式枚举特征:把模型必须保持的每项属性列出来,而不是写"保持角色一致"。
示例提示词(全模态参考):
@image 1: character reference — preserve exactly: silver-white
bob-cut hair, violet eyes, sharp jawline, black turtleneck, silver
pendant necklace.
A character with @image 1's appearance stands in a neon-lit
cyberpunk alley. Rain falls. Reflections on wet ground.
0–4s: Medium shot, frontal. She looks directly at camera. Rain
hits her shoulders. Neon signs reflect on her face — alternating
pink and blue. Camera slowly pushes in.
4–7s: Cut to profile view. She raises her right hand. A
holographic interface appears from her palm. Camera tracks from
her face to the hologram.
7–10s: Over-the-shoulder shot looking at the hologram. Data
scrolls across it. Camera slowly rotates around to face her
through the translucent display.
BGM: ambient synth, low and pulsing
Ambient: rain, neon buzz, distant sirens5. 产品与电商营销
从产品图出发,用提示词把静态照片变成 360 度展示、材质特写、人体工学演示、真实场景植入。
根据对运镜控制的需要选模式:单张产品图做首帧适合简单的旋转展示,多角度产品图做 reference 适合复杂的产品揭幕。
示例提示词(首帧模式):
First frame: product photo — a wireless ergonomic mouse on a
white surface, viewed from a 45-degree angle.
The mouse sits on a clean white desk. The camera showcases the
product from multiple angles.
0–3s: Camera slowly orbits the mouse clockwise, 90 degrees.
Clean white background. Soft diffused light from above.
3–5s: Close-up of the side grip texture. Camera pushes in until
the surface pattern fills the frame. Sound: soft click of the side
button.
5–8s: Pull back. A hand enters from the right and naturally grips
the mouse. Medium shot showing the ergonomic curve against the
palm. The desk surface changes to dark walnut.
8–10s: Wide shot of a full desk setup — monitor, keyboard, and
the mouse in use. The hand clicks and scrolls. Warm office
lighting, shallow depth of field on the mouse.
BGM: N/A
Ambient: quiet office, keyboard tapping, mouse click6. 角色、物体与场景编辑
绿幕抠除、背景替换、台词替换。提示词要说清什么保留、什么替换。
背景替换的核心是光线匹配——如果角色身上主光从左边打来,新背景的光源也要在左边。光线方向不一致是最常见的穿帮。
台词替换的做法是把目标音频作为 reference 上传,提示词要求模型调整角色的口型和表演来匹配新音频。
示例提示词(全模态参考,背景替换):
@image 1: character on green screen — preserve the character's
appearance, pose, and left-side key lighting exactly.
@image 2: new background — a Japanese garden in autumn, warm
afternoon light from the left.
Place the character from @image 1 into the scene from @image 2.
Match the lighting direction — key light from the left in both.
Remove the green screen entirely.
0–5s: Medium shot. Character stands in the garden. Maple leaves
drift down around her. Gentle camera sway, handheld feel. Sound:
birdsong, light wind through leaves.
5–10s: Slow push-in to close-up. Leaves continue falling. One
leaf lands on her shoulder. She brushes it off. The lighting on
her face matches the garden's warm afternoon tones.
BGM: traditional koto, very quiet
Ambient: garden birds, wind, rustling leaves
Dialogue: none常见错误
用氛围词代替动作。"充满电影感"不是指令。"镜头在 3 秒内缓慢推进,角色垂下目光"才是。
上传了素材但不标注角色。 传了三张图片,没说哪张是角色、哪张是风格、哪张是背景。模型随机取用。
参考素材互相打架。 两张不同人的脸。模型取平均,出来的人两边都不像。
写了一堆画质标签。 "4K""hyperrealistic""8K ultra HD"这些不映射到模型参数上,纯粹浪费提示词空间。
一次生成塞太多东西。 10 秒片子里塞五次场景切换、三个角色、一段对白。简化到每次生成一个清晰的戏剧节拍,生成完再剪到一起。
不管结束帧。 如果这条片子要跟其他片子剪在一起,写清楚最后一帧停在什么画面上。结尾没规划好,后期剪辑会很痛苦。
延伸阅读
- MiniMax H3 API 指南——接口对接、参数和异步轮询
- 海螺3 开源权重下载——Hugging Face 和魔搭社区下载、ComfyUI 部署、许可证限制
- 海螺3 是什么——能力边界、局限和定价概览
- 官方 API 文档
- 海螺AI——网页端生成
- MiniMax Hub——社区提示词和示例
- Hugging Face 模型页
- 魔搭社区镜像
上面的模板覆盖了大部分制作场景。从一个镜头开始,用四项检查验证成片——角色一致性、动作完整度、音画同步、剪辑点——然后每次只改一个变量重新生成。写 AI 视频提示词更像写分镜表,不像写作文。
更多文章

MiniMax H3 vs Seedance 2.0:四组实测对比
MiniMax H3 和 Seedance 2.0 到底谁更强?我们用游戏 UI 动效、2D+3D 融合、爆款复刻、一镜到底四组测试实测对比,附评分和花费。


海螺3 是什么?能做什么、效果怎么样、多少钱
海螺3(MiniMax H3)是能同时生成画面和声音的 AI 视频模型。这篇讲清它能出什么样的片子、擅长和不擅长什么、怎么开始用、价格多少。


海螺3 开源权重下载:Hugging Face、魔搭社区与 ComfyUI 部署指南
MiniMax H3 开源权重怎么下载?330 亿参数、Hugging Face 与魔搭社区下载、ComfyUI 部署、许可证地域限制、API 定价——这篇一次讲清。
