MiniMax H3 Max Guide: What It Is, Settings & Prompts
2026/08/30

MiniMax H3 Max Guide: What It Is, Settings & Prompts

Learn what MiniMax H3 Max does, how its launch free tier works, when to use text-to-video or image-to-video, and which settings and prompts to test first.

If you searched for MiniMax H3 Max, you probably want four answers before anything else: Is it different from MiniMax H3? Can you try it for free? What settings should you use? And is the speed claim useful outside a demo?

Here is the short version. According to the August 30 launch materials, H3 Max is a post-trained MiniMax H3 variant aimed at fast 768p text-to-video and image-to-video with native audio. The published test reported about three seconds of backend inference for a five-second clip; that is not a guarantee of the complete click-to-result wait time. The launch web tier also listed a small daily free allowance.

This guide shows how to use it as a creator, not how the launch was marketed.

H3 Max at a glance

QuestionPractical answer
What does it make?768p AI video with native audio
Which inputs work?Text-to-video and image-to-video
How long is one clip?The launch web demo was checked at 5 seconds; verify longer durations on your access point
How fast is it?About 3 seconds of inference for a 5-second clip under the published test
Can I try it free?On August 30, the launch page listed five 5-second generations per day without sign-in
Is it the same as MiniMax H3?No. It uses H3 as its base but targets faster 768p iteration
Is H3 Max Live included?No. Live is a separate experimental continuous-streaming model

The free allowance and service limits can change. The values above were publisher-stated conditions checked on August 30, 2026, not an independent benchmark or a permanent quota. Always read the current control beside the Generate button before planning a batch.

What is MiniMax H3 Max?

H3 Max was presented as starting from the open-weight MiniMax H3 base and improving prompt adherence, aesthetics and serving speed. Think of it as an H3 variant aimed at fast creative iteration rather than a replacement for every H3 workflow.

The practical differences are easier to understand than the model history:

  • H3 Max: fast 768p text-to-video and image-to-video, native audio, short feedback loop;
  • base MiniMax H3: official 768P and 2K output, broader multimodal reference workflows and open-weight deployment;
  • H3 Max Live: experimental stateful streaming that keeps context across scenes.

If your goal is to test ten versions of a social clip, H3 Max is the relevant one. If you need a 2K master, multiple video or audio references, or local weights, start with base H3 instead.

Who should use H3 Max—and who probably should not

H3 Max is most useful when you need to compare controlled variants quickly: hooks, camera moves, timing or product-image animation. An agency can show a client moving storyboards instead of debating timing from still frames, while a product team can compare camera directions before paying for a final shoot or a higher-resolution render.

It is a particularly good fit for:

  • short-form creators who need several openings, transitions or visual hooks for the same idea;
  • storyboard and previsualization work where timing and camera movement matter more than final pixels;
  • product-motion studies built from an approved product image;
  • dialogue blocking where you need to see whether one short line fits the shot;
  • small teams that benefit from immediate feedback instead of overnight render queues.

It is a weaker fit when the first requirement is exact text, exact packaging, 2K delivery, a large reference pack or a long scene with several edits. Fast generation does not remove those constraints. Use H3 Max to answer one creative question at a time, then move the approved direction into the workflow that can meet the final-delivery requirement.

Your real questionGood H3 Max test
Does this hook read in three seconds?Generate three five-second openings with the same subject
Which camera move sells the product best?Hold the image and action constant; change only the camera
Can this line fit naturally?Use one speaker, one sentence and a locked shot length
Will the client approve the composition?Animate the approved storyboard or product frame
Can it make the final 2K master?No—use the test to choose a direction, not to pretend the requirement disappeared

How to use H3 Max without code

1. Choose text-to-video or image-to-video

Use text-to-video when the composition can be invented from scratch. It is good for ideation, establishing shots and visual directions where you do not already have approved artwork.

Use image-to-video when the first frame matters: a product photo, character design, storyboard panel or approved composition. The input image does not lock every later frame, but it gives the model a much stronger starting point than text alone.

2. Start with five seconds

A five-second test is the cheapest way to answer the important questions:

  • Does the subject look right?
  • Does the camera move in the requested direction?
  • Does the main action finish?
  • Is the dialogue short enough?
  • Does the generated audio belong to the scene?

Do not begin with a longer duration just because an access point offers it. Longer prompts expose more failure points and may consume more quota or credits than a five-second diagnostic, so validate the shot at five seconds first.

3. Use 16:9 unless the destination is already known

Choose the final platform ratio before you write the composition:

DestinationRatio to test firstComposition note
YouTube or website hero16:9Keep the subject away from the extreme edges
TikTok, Reels or Shorts9:16Ask for a medium or full-body vertical composition
Instagram feed1:1Limit lateral camera travel

Changing aspect ratio after you approve a shot is effectively a new composition. Test the real delivery ratio early.

4. Write one shot, not a whole commercial

A one-shot brief is easier to diagnose than a paragraph containing several edits. A reliable prompt has five parts:

[subject and fixed details] + [one visible action] + [location and lighting] + [camera move] + [dialogue, ambience and constraints]

Example:

A ceramic artist in a clay-stained navy apron lifts a newly glazed cup from the workbench and turns it once toward the window light. Quiet workshop at blue hour, shelves of unfinished pottery behind her. Slow 50mm push-in, shallow focus, natural hand movement. Sound: soft room tone, distant rain and one ceramic tap. No captions, no logos, no cut.

This gives the model one subject, one action and one camera plan. If you need three shots, generate three clips. Do not hide three edits inside one sentence and hope the model invents the same timing you had in mind.

Three H3 Max prompt recipes to copy and adapt

The five-part formula becomes easier to use when you start from the job rather than from a list of cinematic adjectives.

Recipe 1: a five-second product-motion test

Use image-to-video and treat the uploaded packshot as the composition and packaging reference, while assuming that text and geometry still need frame-by-frame review.

The same matte silver travel bottle remains centred on the stone plinth. A narrow band of morning light moves slowly from left to right across the bottle while two condensation drops travel down the front. Camera makes a gentle 10 cm push-in at product height, 50mm lens, no tilt. Keep the cap shape, bottle proportions, printed label and silver colour unchanged. Sound: quiet room tone and one soft glass tap. No hands, no new props, no text change, no rotation.

Why it is written this way: the prompt names the approved elements, limits motion to two visible changes and gives the camera a measurable path. It cannot guarantee exact packaging, so inspect the label, cap and proportions before using the result. The common mistake is asking for “a premium commercial” and leaving the product constraints unstated.

Recipe 2: one character, one short line

A tired night-shift baker in a flour-dusted green apron leans against the stainless-steel counter and looks toward the first light outside. She exhales, then says quietly, “We made it.” Warm oven light behind her, cool dawn at the window. Locked medium close-up with a very slow push-in. Preserve her short dark curls, green apron and the scar above her left eyebrow. Sound: low refrigerator hum, distant delivery van, no music. No cut, no subtitles, no second speaker.

Why it works: the line is short enough for the duration, the speaker is unambiguous and three visible identity anchors protect the character better than a paragraph of biography.

Recipe 3: an atmospheric establishing shot

A small coastal train crosses a weathered steel bridge during a summer storm, warm carriage windows reflected in the dark water below. One sheet of rain moves across the valley as lightning briefly reveals the distant cliffs. Wide 24mm aerial tracking shot, camera follows parallel to the train at constant speed. Sound: rain, rail rhythm and one distant thunder roll. Natural scale, no speed ramp, no cut, no text.

Why it works: the train supplies a clear motion line, the weather supplies one timed visual change and the camera relationship stays stable. Avoid adding a second location or asking the shot to move from exterior to inside a carriage in five seconds.

Best H3 Max settings for a first test

Start with this baseline:

SettingRecommended first testWhy
ModeText-to-video for ideas; image-to-video for approved compositionMatches the job instead of forcing one workflow
Resolution768PCurrent H3 Max target
Duration5 secondsMatches the checked launch demo and gives a short diagnostic
Aspect ratioFinal delivery ratioAvoids approving a composition that will later be cropped
Prompt expansionLeave at the default for the baseline; test separately if exposedPrevents an extra variable from obscuring the first comparison
AudioDescribe only sounds that should be heardPrevents generic music from fighting the scene

If your access point exposes a longer duration, extend only after the five-second version works. Add a timeline rather than extra adjectives:

0–5s: she lifts the cup and checks the glaze.
5–10s: she turns toward the window and the surface catches the blue light.
10–15s: she sets it beside the finished set and smiles, camera settles.

Each interval needs a visible change. Empty time often becomes slow motion, repeated movement or an unnecessary camera drift.

A five-run test plan that teaches you something

Five random prompts produce five unrelated opinions. A controlled five-run sequence tells you which part of the brief caused the improvement.

  1. Baseline: generate the simplest five-second version with one action and one camera move.
  2. Camera test: keep every word the same except the lens and camera path.
  3. Action test: restore the original camera and change only the subject's action or timing.
  4. Reference test: if identity or composition drifted, switch to image-to-video using the best opening frame.
  5. Delivery test: if the destination was unknown during ideation, switch to the real aspect ratio now; otherwise keep the final ratio from the baseline. Extend duration only if the access point supports it.

Keep a small test log instead of trusting memory:

RunVariable changedKeep / rejectReason
1BaselineSubject, action, camera, audio
2Camera onlyDid the product or character read better?
3Action onlyDid the beat finish inside five seconds?
4Input imageDid identity and composition improve?
5Ratio or durationDid the approved idea survive delivery settings?

Change one variable per run so it is easier to identify what improved the result. This reduces interference but does not prove causation because generation remains stochastic. If a fixed seed is available, keep it constant; if it is not, repeat important variants two or three times before choosing a winner.

How to write motion constraints for image-to-video

The earlier mode choice decides whether an input image is useful. Once you choose image-to-video, the next job is to separate what may move from what must remain recognisable. Describe what changes after the input frame:

The camera slides ten centimetres to the right while the bottle remains centred. Condensation beads roll slowly down the glass; the backlight brightens and the reflection travels across the label. Keep package geometry, label text and cap colour unchanged. No rotation, no new objects.

The phrase “keep everything the same” is too vague. Name the parts that matter.

H3 Max versus MiniMax H3

NeedBetter starting point
Generate many ideas quicklyH3 Max
Short dialogue or social clips at 768pH3 Max
2K final outputMiniMax H3
Multiple image, video or audio referencesMiniMax H3
Open-weight or local workflowMiniMax H3
Continuous interactive broadcastH3 Max Live, when production access is available

A sensible production workflow can use both. Explore with H3 Max, keep the prompt and approved frame, then decide whether the selected shot needs a higher-resolution H3 render.

Common H3 Max problems and quick fixes

The clip ignores the second half of the prompt

Cut the brief to one action. Put the visible event before style words. If two beats are essential, assign each a time range.

The character changes while moving

Use image-to-video and repeat two or three identity anchors: hair shape, one garment and one distinctive accessory. A long biography does less than three visible details.

Dialogue sounds rushed

Shorten the line and read it aloud at a natural speaking rate. Five seconds usually fits one short sentence, not a paragraph. Bind it to one speaker and describe the tone once.

The output feels slow despite fast generation

Model speed and shot pacing are different. Replace vague verbs such as “moves gracefully” with a physical action and a camera path. If the subject has nothing to finish, the model fills time by slowing down.

The result is fast but not delivery-ready

Use H3 Max for iteration, then judge the selected shot at final size. 768p can be enough for social and concept work; it is not automatically the right master for a large display or a crop-heavy edit.

H3 Max questions people ask before the first generation

Is the “three seconds for five seconds of video” claim the total waiting time?

Treat it as an inference-speed result, not a promise that every click finishes in three wall-clock seconds. Upload time, queueing, safety checks, audio processing and delivery can add time around the model run. Compare complete click-to-result time on the access point you actually use.

Does native audio mean I can skip sound design?

No. Native audio is useful for dialogue timing, ambience and concept review. Final work still needs a listen for speech clarity, unwanted music, abrupt endings and consistency across separately generated clips.

Can H3 Max keep the same character across several clips?

It can reuse visible anchors, especially with image-to-video, but separate jobs do not become one continuous memory automatically. Save the approved first frame, repeat the same visible identity details and expect to repair continuity in the edit.

Should I write prompts in English?

Use the language in which you can describe the shot most precisely. If a specific camera or lighting term behaves inconsistently, keep that production term in English and write the rest naturally. A clean mixed-language prompt is better than a vague full-English translation.

Can I use generated clips commercially?

Model quality and usage rights are separate questions. Check the current service terms, confirm you own or may use every uploaded image, and review the output for protected logos, faces, music and other material you did not ask for. A fast render does not clear rights for you.

Is H3 Max worth using?

Use it when fast feedback changes how many ideas you can test. It is especially useful for social clips, visual development, dialogue blocking, product-motion studies and previsualization.

Skip it when your first requirement is 2K delivery, a large multimodal reference pack, local deployment or a persistent live session. Those needs point to base MiniMax H3 or H3 Max Live instead.

For prompt structures you can adapt, browse the MiniMax H3 prompt library. The models share a base, so the subject-action-camera-sound structure transfers well, but always validate the final result on H3 Max itself.

If your interest is the continuous version rather than ordinary short clips, read H3 Max Live for livestream creators or build a scene cue with the H3 Max Live Prompt Director.

Product limits and free-access details checked August 30, 2026. They are service conditions, not permanent model guarantees.

Free to try

Generate your first video with MiniMax H3 — right now

Create videos from text or a reference image, with ready-to-use prompt examples to help you get started. No downloads — just open it in your browser.