Five weeks have trickled by since MiniMax set free MiniMax H3. The model shook the whole creative tech crowd. ComfyUI now rolls out its own set of ready-to-go templates. Hugging Face stores the weights. Those graphics card services already offer simple access. Yet, something important never showed up. The world got flooded with user manuals, but prompt advice always circled back to the same huge pile of official examples. Those guides only repeat what MiniMax already gave. So, beginners probably feel seen. Hardcore tinkerers? No help at all. four native H3 templates Hugging Face One of them says so in its own title

I have now been consumed by MiniMax. The more I test, the more capabilities I have found. Two weeks ago, I would run a four tool flow – a Picture Generator, a Picture Editor, a Video Creator, and an Audio Addition. Each step is a different ComfyUI run. The final result was a 15 second clip. For a six minute music video, about 50 finished clips. I have about an 80% failure rate so figure 300 clips total. It took me about a month of weekend work to create a 6 minute finished video.
With MiniMax I am finding that I can use MiniMax to edit my picture, and it has much better Audi than LTX, and the Video Engine continues to amaze me in allowing me to add creative elements. I honestly think I can prompt – “The standing man gets on an Elephant and Rides to Paris” and it will do it. (Testing now: Results – without any reference H3 created an elephant and my character properly got on and rode it. No trace of Paris.) Counting failed generations about 300 clips total. week, my own computers and those rental clouds spun up H3. Some tips from the official playbook do what they promise. Some give absolutely nothing. Unspoken rules hide beneath the surface. These patterns came from mistakes. Skimming hardware specs revealed none of them. Anyone frustrated over fake tunes, switched dialogue, or a scene suddenly forgetting a reference might want to look below. The answers probably live here.
H3 Prompting
H3 prompts are a lot different than Stable Diffusion, Wan and LTX.
Many creative tools employ CLIP. That mysterious tool grew from Contrastive Language-Image Pre-training. CLIP may barely notice full sentences. Around seventy words slip through at a time, and the system sniffs around for meaning. Sentence order? Minor. Grammar? Almost invisible. This explains why everyone tosses tags with wild commas. Users keep repeating ideas because CLIP just shrugs at them.
MiniMax H3 rewrites the rules. A large language engine handles the prompts. The tech behind chatbots, called LLMs, fills the job. Qwen3-VL-32B became the brain for H3. Whole sentences matter now. Order matters. Memories from earlier phrases stick. So H3 might actually follow requests, not just notice stray words.
Everything you knew about prompting for Stable Diffusion might just fall apart.
Attention Weight does nothing
Those who came from Diffusion tools love to boost settings with weights. That familiar command, (very slowly:4.0), pushes concepts further. H3 reads each symbol literally. Extra punctuation or numbers get ignored. Changes never appear.
Use time instead. H3 might actually grasp speed. “Each wave takes about two seconds” may work. “Very slowly” might have no effect at all. H3 allows adjective modifiers – very slow, extremely slow. Beware of metaphors – glacially slow may or may not work.
Long prompts work, but keep the grammar plain
I have seen examples where other people use long and texty prompts. I have not tested that. My current opinion is that 90% of their prompt is dead weight. A long prompt takes longer to maintain and adjust than a shorter one. The key is to discover which words are holding weight, and to use them carefully.
In my testing, a flat, direct style gives stronger results than gentle English. And non-grammatical prompts “broken English” work better than proper grammar. If you say, “Max do gawk,” you possibly get better action than, “Max performs a gawk.”
When I first started video generation and I think with Wan, I would use Google Translate to convert my prompts to Chinese (AI Engines speak Chinese) and back. And also I’m somewhat familiar with Asian languages. So plurals collapse, “a” or “the” have no meaning, and adjectives are simple. There’s no real difference between “The man stands up and walks out the door” and the Chinese equivalent, but “The man rises and leaves the room” may not prompt so well. It has the same meaning, but the words that work are more mechanical and more CEFR A1 English rather than B2.
H3 allows you to define actions based on keywords
My most exciting find in H3, that no other engine has, is that there appears to be some programming capability. It is not perfect, but I get some amazing results.
Definitions bind, and they take arguments
Inside a prompt, you might teach H3 a new word and call on it later.
A gawk means a person throws an object outward from their body into the air and out of view.
There are three people in this picture, Max, Mary and Monk.
Max do gawk with yellow ball. then Mary do gawk with green chicken.
then Monk do gawk with blue cat. then Max do reverse gawk with blue cat.
Here I am creating a definition in my prompt and then using it with different characters. With H3 – It followed my rule. Right thing, right person. Not always, but most of the time. The gawk idea acts like a reusable shape with a gap you can fill each time.
It fails to overcomplicate a defined term. In my sample above “reverse gawk” did not work at all. The character did a gawk in the vide, and ignored “reverse.” But I still find this a highly exciting test.
In real life – I think this could allow me to make more human motion. For instance a “queenwave” means to wave your hand gently side by side, without tilting the fingers. “Jesse does a queenwave, then Mary does a queenwave”
Modifiers do not compose
Begin with one clear line to set up who is who. “Two people appear. Max stands to the left, Mary to the right.” After that line, i just use names. I have seen prompting for Subject 1 or (S1) and I have tested that. It did not work as well for me as naming my characters and using the names. There is not strict adherence to the names particularly for characters that are not fully on screen.
Bind names to positions once, then never again
Never link positions such as “left” or “right” in every line. When characters switch spots, every mention flips. Miss one and the entire story breaks. One setup line, fixed in one place, saves the day. Twenty random references guarantee trouble.
H3 appears to have character memory so if my right and left person cross – I can continue to prompt them as Max and Mary.
How to Get the Audio You Asked For and Nothing Else
Two keys control the prompt’s audio. One key may steer the overall_soundscape. The other key operates non_diegetic_music – music that floats in but comes from no object in the scene.
N/A is the off switch, and it is literal
You can switch off either key by writing N/A. Do not invent new terms. Only N/A works, straight from MiniMax’s own prompt notes. Ask for “None.” and H3 will probably guess. “No music please” does not work. “none” does not work.
Say nothing about music and H3 supplies a soundtrack. Every single time. Silence in your lines means nothing. If your aim is no music, non_diegetic_music: N/A becomes the only way. MiniMax’s own system prompt
For now I put this on the top of every prompt:
overall_soundscape: N/A
non_diegetic_music: N/A
integrated_multimodal_description:
[Shot 1] <Picture 1> is fully referenced.
What follows is the content I would normally prompt for a Wan or LTX video. With this start, I’m telling it to not generate background music, and also do not generate background noise that I do not specify, and to use Picture 1 as the start.
In addition, I have found that generally speaking I do not have to introduce the picture as I do with Wan or LTX. For those I would start with – there are 3 people in this picture, a right person who is a man, a middle person who is a woman and a right person who is a woman monk with shorn hair. Now, I sometimes name the figures, but it is not really needed unless I use the names lower down for dialog. This is just what I found and may not be accurate.
Also about 10% of the time I DO get diegenic music, and about 50% of the time I DO get background sounds – despite my prompt.
Quote the dialogue or you get babble
Leave the lines unwritten and H3 does not go quiet. It invents speech-shaped audio in no identifiable language. It sounds vaguely Romance. It is unusable. If you want words, type the words.
How to Make the Right Character Speak
Mouth movement seems to be the trickiest part of H3. The voice of one character might just leak to someone else, which has puzzled plenty of users. A known problem with the model sits at the heart of this struggle, not simply bad instructions. Discussions have popped up among developers searching for a reliable fix. open ComfyUI issue
Is the d tag required?
H3 documentation says <d> marks speech. Inside the dialogue tag, place just the language marker and the words. Leave everything else outside. Speaker names, clues about how their voices sound, and even their delivery style must stand apart from the dialogue tag. Anything else might confuse the model.
Here is an example that probably helps:
Max, older and deep-voiced, says: <d>[English] I told you it was cold.</d>
Mary responds: <d>[English] Watch me.</d>
The scene becomes clear. Voices connect to the right faces.
Now as a director, I’m highly resistant to this style. So what I would do today is this:
Max: "I told you it was cold."
Mary: "No it is not!"
and for me this seems to work ok – at least as often as the <d> framing. Right now I’m also placing the voice timbre queue in the first use of the voice:
"Max in a high squeaky voice: "I told you it was cold".
Name who stays silent, or change your script
H3 has a very annoying habit of having both people speak the same phrase at the same time.
Right now I’m testing this:
Max: "I told you I was cold" - Mary is silent.
Mary "Watch me" - Max is silent.
Sometimes works.
Sometimes in frustration I re-script a clip to match whoever H3 is favoring in the script. I don’t fight it, I join it.
Off-screen speech needs the exact phrase
When H3 has moved camera focus on one character, the other may be off screen or only partially on screen. If their mouth is not visible, H3 seems reluctant to have them speak. Assigning dialog to the off screen character does not work. So I prompt it separately.
Voice from off screen "I mean it, it is very cold".
It seems to be based on whether the character’s lips are visible.
Write these words exactly—off-screen voice says:—and the line probably works perfectly. Use “offscreen” rather than “off-screen” and it won’t work at all. The system is still internally highly dependent on the words, possible to interpret them in Chinese or through some internal dictionary.
Voice timbre works, but placement decides it
Voice control seemed impossible until I figured out the placement. I tried it on the top or independently. But two things seem to work for me. One is in the character intro:
The left man is Max who is 19 years old and has a high squeaky voice.
Or I can put it in the prompt.
Mary (in a low elegant tone): "It is not cold at all"
I run a lot of side-tests to check the placment. Sometimes works, sometimes not.
Spoken language comes from the picture
The model might guess the speaking language from a reference image, not the prompt. Pictures of people from East Asia, for example, somehow produced Mandarin even when I wrote lines in English. Language can slip away from your plan that quickly.
Anchor the language in your script. Adding a line such as “They speak English” usually keeps H3 on task. Hinting at a nationality sometimes works, but may also change the character’s look—so that trick remains unpredictable.

Timing and the Things That Fight You
H3 fills the duration you give it
H3 will probably not slow down to fill the moments you want. Give H3 more seconds than action beats and it might just invent extra scenes or repeat content. The best results come from matching your video’s timing to the number of actions you planned.
One physical event per clip
Physics looks believable for single actions. A basic throw always works. Try combining a throw with several bounces and H3 probably fails, in every phrasing I tried. Breaking big actions into smaller moments gives much better results.
Motion direction is random unless you state it
Sometimes, actions run in reverse. Once, the ball traveled backwards into a hand instead of away. The model might have learned this as normal from reversed footage in training, so no default direction exists. Describing direction clearly—“away from the body, forward into the distance”—locks the motion in. Later references keep the action moving as planned.
H3 takes hints from what you feed it. I have seen raised arms in the first frame become a throw before prompts even begin. Picking a frame where everyone looks relaxed and neutral may prevent these surprise movements. The starting moment shapes everything that follows.
The start frame’s implied motion happens anyway
H3 reads intent out of your source image. A raised arm in my still became a throw, before the prompted action even started. Choose start frames with neutral posture unless you want that motion.
Shot markers cause the cuts you are avoiding
One time, I tried to use [Shot 2] and [Shot 3] for splitting story moments. That was a mistake. Those tools tell the system when to cut scenes. Unwanted scene breaks kept popping up when I just wanted smooth flow. Now, every action lives together in a single Shot 1 for me.
Failed Generation Due To Size Change
At one point I had 20 clips all tested in small size, and I queued them to generate production size. All 20 clips failed. About 5 seconds in – H3 decided to change the entire background and characters. But the exact prompts worked perfectly in my test to the small size.
And then I found this – Clips tend to change when they hit a certain pixel limit. The larger your image size the sooner that limit is. I have three models in my workflow:
MiniMax H3 - Turbo on - 4 Step Lora
MiniMax H3 - Turbo on - 8 Step Lora
MiniMax H3 - Turbo off - 20 Step
What I found is that I can test small videos quickly using 4-Step. But for production I have to move up to 8 Step or 20 Step to maintain the integrity and quality through the entire 15 second video.
I have not found at all that integrity is lost if I move to 20 or 25 seconds. The motion never repeated. But on my machine, the video production runs out of memory at the longer sizes, and so there is a physical limit to the size/length I can test. For me, 796p at 15 seconds is Max on my PC, and I can go to 20 seconds on an RTX 6000 RunPod, but that’s the extent I have worked with.
If your background is shifting – if sound is degraded – move to the next higher Turbo/Step model and see if it improves.
The Short Version
When writing, go long and use simple language. Describe every action just once, then call it by its chosen name later. Put all names and their spots right at the beginning. Set both audio options to N/A, unless you really want some surprises. Use quotes for each line people speak. Always choose the d tag. At a character’s first words, mention the voice. Clearly say the language. Keep one movement or action for each short clip. Try to match every clip’s length to your dramatic beats. Throw out the shot markers.
None of these secret moves appear in the official instructions. Bad drafts taught me each one. Anyone planning to share work made with H3 should check the terms carefully. People in the United States and the European Union might find a real problem hidden in the rules.


















