MiniMax H3长视频生成器
H3 单次只出 4 到 15 秒,更长的都是分段接起来的。这个规划器帮你拆好段、为每一段写对提示词——包括几乎所有人都写错的那句续接指令——并在你花钱之前先把账算清楚。
先给结论
MiniMax H3 能做 60 秒视频吗?
单次生成做不到。H3 线上接口只收 4 到 15 秒,且必须是整秒,这是硬上限。60 秒或 120 秒的 H3 视频,永远是若干段首尾相接——每一段都从上一段的最后一帧开始。这套做法叫视频续接,它确实能用,但它是一串独立推理,不是一次长渲染。
- 单次生成:4–15 秒,自带同步原生音频。
- 60 秒 = 4 段 × 15 秒;120 秒 = 8 段 × 15 秒。
- 第一段之后的每一段,都是以上一段末帧为输入图的图生视频任务。
本页负责把这条链规划好、把提示词写好。跑分段和合并文件仍然要你自己一段一段来——与其卖你一个还不存在的按钮,不如把话说清楚。
规划器
把总时长拆成 H3 真收得下的分段
设定总时长,描述人物和场景,再给每一段一个动作节拍。下面的提示词可以直接粘——英文,因为那才是 H3 训练时用的格式。
每段一个节拍
这一段里发生了什么变化。写动作,别写概括——没事可做的一段,会飘或者会循环。
你的分段
4 段 · 60 秒 · 合计约 804 积分
按顺序跑。每一段跑完,取它的最后一帧当作下一段的输入图。
第 1 段 · 0–15 秒 · 共 15 秒
For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced. integrated_multimodal_description: [Shot 1] Live-action, cinematic, a courier in a red rain jacket with a frayed left cuff, mid-twenties, short black hair, in a neon-lit night market street, wet asphalt, stalls still open, matching the position, framing, and lighting established by <Picture 1>. she pushes off the kerb and threads between two stalls, checking the address on her phone. overall_soundscape: Steady rain on canvas awnings, tyres through standing water, market chatter fading in and out, a tram bell in the distance. non_diegetic_music: N/A
第 2 段 · 15–30 秒 · 共 15 秒
For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced. integrated_multimodal_description: [Shot 1] Live-action, cinematic, a courier in a red rain jacket with a frayed left cuff, mid-twenties, short black hair, in a neon-lit night market street, wet asphalt, stalls still open, continuing without interruption from the position, framing, motion direction, and lighting established by <Picture 1>. she slows at a crossing as a tram passes, rain beading on the jacket. The identity, clothing, colour palette, key props, and lens character stay identical to <Picture 1>; no cut, no restart, no change of shot scale at the opening. overall_soundscape: Steady rain on canvas awnings, tyres through standing water, market chatter fading in and out, a tram bell in the distance. non_diegetic_music: N/A
第 3 段 · 30–45 秒 · 共 15 秒
For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced. integrated_multimodal_description: [Shot 1] Live-action, cinematic, a courier in a red rain jacket with a frayed left cuff, mid-twenties, short black hair, in a neon-lit night market street, wet asphalt, stalls still open, continuing without interruption from the position, framing, motion direction, and lighting established by <Picture 1>. she lifts the parcel from the crate and scans the shopfront numbers. The identity, clothing, colour palette, key props, and lens character stay identical to <Picture 1>; no cut, no restart, no change of shot scale at the opening. overall_soundscape: Steady rain on canvas awnings, tyres through standing water, market chatter fading in and out, a tram bell in the distance. non_diegetic_music: N/A
第 4 段 · 45–60 秒 · 共 15 秒
For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced. integrated_multimodal_description: [Shot 1] Live-action, cinematic, a courier in a red rain jacket with a frayed left cuff, mid-twenties, short black hair, in a neon-lit night market street, wet asphalt, stalls still open, continuing without interruption from the position, framing, motion direction, and lighting established by <Picture 1>. she hands the parcel over at a lit doorway and steps back into the rain. The identity, clothing, colour palette, key props, and lens character stay identical to <Picture 1>; no cut, no restart, no change of shot scale at the opening. overall_soundscape: Steady rain on canvas awnings, tyres through standing water, market chatter fading in and out, a tram bell in the distance. non_diegetic_music: N/A
会打开图生视频生成器并带上这段提示词。生成需要登录和积分。
花积分之前
没有需要提醒的——这条链已经尽可能连续了。
原理
"视频续接"到底在做什么
H3 接口上没有 extend 参数。续接是一种约定,不是一个端点:你把上一段的最后一帧——就一张静止图——交给下一次生成,让它接着往下演。
- 1
生成第 1 段
文生视频,或从参考图做图生视频。这一段定下后面所有段的身份、色调和镜头质感。
- 2
取它的最后一帧
把成片停在最末尾导出那一帧,或者在任何剪辑软件里抽出来。这张图就是下一段的输入图。
- 3
下一段按 I2VA 跑
图生视频,输入图是刚才那一帧,提示词用规划器给的续接版。不是首尾帧——此刻还没有第二张图。
- 4
重复,然后合并
循环到够时长为止,再把这些片段在剪辑软件里接起来。合并之前,它们是彼此独立的 MP4。
能跨过接缝的东西
- 切点上的构图和取景——那本来就是同一帧。
- 你在提示词里一字不差重复的部分:主体、服装、场景、风格词。
- 大范围的色彩和光线,前提是你交出去的那一帧曝光正常。
每到接缝就重置的东西
- 运动。静止帧不携带速度,模型只能重新推断动作,快动作会在切点上明显一顿。
- 音频。每段各生成各的,环境音和配乐都会重启,除非你把它们描述得一模一样。
- 画外的一切。没出现在那一帧里的东西,对下一段来说就等于不存在。
这是大多数长视频教程写错的一条:只有第 1 段——或者必须落到指定尾帧的那一段——才是 FL2VA。所有续接段都是 I2VA,而 I2VA 用的是另一句指令。只给一张图却套 FL2VA 那句话,这张图会被悄悄忽略。
预期
一条链能撑多久
每一段能看到的只有一帧历史,加上你的提示词,再无其他。这一条事实,就足以推出什么会坏、大概什么时候坏。
| 链条深度 | 通常还稳 | 开始出问题 |
|---|---|---|
| 第 1–2 段 | 身份、服装、色调、场景。 | 第一个接缝上的动作节奏,动作越快越明显。 |
| 第 3–4 段 | 有锚点重复的前提下,身份和场景仍在。 | 色温开始偏;次要道具变形。 |
| 第 5–6 段 | 大场景和被描述过的锚点。 | 面部变软、比例改写;服装细节自己重写。 |
| 第 7–8 段 | 风格词和大致设定。 | 把第 1 段和这段并排一看,累积偏移通常已经很直观。 |
这些是由机制推出、并在社区反馈里反复出现的失效模式,不是我们跑过的 benchmark 实测值。等我们发数字的时候,会一并说明是怎么测的。
成本
在这里做一条长 H3 视频要多少钱
H3 按生成秒数计费,所以一条链的成本就等于它的总时长——再加上你重跑的每一段。真正的预算大头是重跑,不是这份规划。
| 成片时长 | 15 秒分段数 | 积分 | 若重跑两段 |
|---|---|---|---|
| 30s | 2 | 402 | 804 |
| 60s | 4 | 804 | 1206 |
| 90s | 6 | 1206 | 1608 |
| 120s | 8 | 1608 | 2010 |
数字取自本站正在用的 H3 计价,所以和生成器实际扣的分一致。生成失败会自动退分;只是你不满意的那次不会。
操作
做一条 60 秒的 H3 视频
- 1
规划这条链
选 60 秒、单段 15 秒,得到四环,然后每段写一个节拍,别让任何一段有空时间要填。
- 2
锁死身份锚点
把主体具体地描述一次,让规划器一字不差地重复进四段提示词。对抗漂移,这是性价比最高的一步。
- 3
用参考图跑第 1 段
打开图生视频生成器,传起始帧,粘第 1 段提示词,按 15 秒生成。
- 4
把后面几段接上
导出每段成片的最后一帧,作为下一段的输入图,再粘那一段的提示词。重复到时长够为止。
- 5
合并片段
在任意剪辑软件里把四个 MP4 接起来。如果切点上能看出动作重启,每个接缝各修掉几帧。
把话说准
被叫成"H3 长视频"的三件不同的事
这三件事经常被混着说,而今天真正能做的只有一件。
| 做法 | 指的是什么 | 现状 |
|---|---|---|
| 单次生成 | 一次推理、一个文件,全程原生音频。 | 可用——上限 15 秒。 |
| 分段续接 | 多次生成,每段都从上一段的末帧开始。 | 今天可用。本页负责规划,跑和合并由你来。 |
| 流式 / 因果生成 | 边播边持续产出,没有固定片段长度。 | 仍在研究阶段。线上 H3 接口没有,本站也不提供。 |
常见问题
MiniMax H3 长视频问答
MiniMax H3 视频最长能做多长?
单次生成 15 秒,最短 4 秒,且必须整秒。文生视频、图生视频、参考图生视频都是这个上限。更长的都得靠分段拼接。
怎么做一条 2 分钟的 MiniMax H3 视频?
按八段 15 秒来规划。第一段从参考图生成,之后每一段都用上一段成片的最后一帧作为输入图,提示词里重复同一句身份锚点。最后把八个文件接起来。要有重跑一到两环的心理准备。
H3 有 extend 或者续写接口吗?
没有。线上接口上没有 extend 参数。续接是靠把上一段的末帧作为下一次生成的输入图来实现的——这也是为什么续接提示词是图生视频的写法,而不是首尾帧的写法。
我能上传自己的视频来续写吗?
本站目前不行。生成器收图片,不收视频。如果你手上有一段想续写的视频,请先导出它的最后一帧,从那张图开始这条链。
为什么接缝上动作会顿一下?
因为静止帧不携带速度。下一段只能从零重新推断运动,快动作就会在切点上像停了一拍。让接缝附近的节奏别太快,合并时每个接缝修掉几帧,都有帮助。
声音能保持连续吗?
默认不能。每段各生成各的音频,除非每一段提示词里都出现同一句声景描述,环境音就会在每个接缝重启——所以规划器会替你重复它。至于配乐,合并之后在剪辑软件里统一配更靠谱。
这是 MiniMax 官方功能吗?
不是。4–15 秒的限制是 MiniMax 定的;分段续接是社区在它之上搭出来的工作流,这个规划器是我们的实现。本站与 MiniMax 无隶属关系。