Five weeks have trickled by since MiniMax set free MiniMax H3. The model shook the whole creative tech crowd. ComfyUI now rolls out its own set of ready-to-go templates. Hugging Face stores the weights. Those graphics card services already offer simple access. Yet, something important never showed up. The world got flooded with user manuals, but prompt advice always circled back to the same huge pile of official examples. Those guides only repeat what MiniMax already gave. So, beginners probably feel seen. Hardcore tinkerers? No help at all. four native H3 templates Hugging Face One of them says so in its own title
I have now been consumed by MiniMax. The more I test, the more capabilities I have found. Two weeks ago, I would run a four tool flow – a Picture Generator, a Picture Editor, a Video Creator, and an Audio Addition. Each step is a different ComfyUI run. The final result was a 15 second clip. For a six minute music video, about 50 finished clips. I have about an 80% failure rate so figure 300 clips total. It took me about a month of weekend work to create a 6 minute finished video.
With MiniMax I am finding that I can use MiniMax to edit my picture, and it has much better Audi than LTX, and the Video Engine continues to amaze me in allowing me to add creative elements. I honestly think I can prompt – “The standing man gets on an Elephant and Rides to Paris” and it will do it. (Testing now: Results – without any reference H3 created an elephant and my character properly got on and rode it. No trace of Paris.) Counting failed generations about 300 clips total. week, my own computers and those rental clouds spun up H3. Some tips from the official playbook do what they promise. Some give absolutely nothing. Unspoken rules hide beneath the surface. These patterns came from mistakes. Skimming hardware specs revealed none of them. Anyone frustrated over fake tunes, switched dialogue, or a scene suddenly forgetting a reference might want to look below. The answers probably live here.
No tip matters until you face the real difference between H3, Stable Diffusion, and Wan.
Many creative tools employ CLIP. That mysterious tool grew from Contrastive Language-Image Pre-training. CLIP may barely notice full sentences. Around seventy words slip through at a time, and the system sniffs around for meaning. Sentence order? Minor. Grammar? Almost invisible. This explains why everyone tosses tags with wild commas. Users keep repeating ideas because CLIP just shrugs at them.
MiniMax H3 rewrites the rules. A large language engine handles the prompts. The tech behind chatbots, called LLMs, fills the job. Qwen3-VL-32B became the brain for H3. Whole sentences matter now. Order matters. Memories from earlier phrases stick. So H3 might actually follow requests, not just notice stray words.
Everything you knew about prompting for Stable Diffusion might just fall apart.
Those who came from Diffusion tools love to boost settings with weights. That familiar command, (very slowly:4.0), pushes concepts further. H3 reads each symbol literally. Extra punctuation or numbers get ignored. Changes never appear.
Use time instead. H3 might actually grasp speed. “Each wave takes about two seconds” may work. “Very slowly” might have no effect at all.
Longer requests will probably help you more. Ranges between several hundred words feel normal. Extra explanation helps H3 stick to your goals. Bonus tags only clutter up your ideas and probably waste words.
Yet, surprises appeared during my testing. A flat, direct style always gave stronger results than gentle English. If you say, “Max do gawk,” you possibly get better action than, “Max performs a gawk.” Qwen models lean toward direct, Chinese-inspired phrasing. Simple order with verb before object wins. Cut out unnecessary words. So — skip articles. Drop tense endings. Write long, but keep each phrase boringly clear.
H3 has internal programmability. It is not perfect, but I get some amazing results.
Inside a prompt, you might teach H3 a new word and call on it later.
I defined “gawk” means someone tosses an object away, hand flying up, until the thing vanishes from sight. Imagine three people in a picture. Their names: Max, Mary, Monk. First, Max performs a gawk with a yellow ball. Then, Mary repeats a gawk with a green chicken. Then comes Monk, who sends a blue cat flying.
With H3 – It followed my rule. Right thing, right person. Not always, but most of the time. The gawk idea acts like a reusable shape with a gap you can fill each time.
CLIP rarely manages this trick. Many platforms, like Wan and LTX, will make you rewrite every step for every person. My usual definitions may stretch for several hundred words. The whole reason for a long prompt lies in defining once, not everywhere.
Stick with then to connect actions. This little word often ends action problems other tricks cannot reach.
It fails to overcomplicate a defined term. I tried “reverse gawk.” H3 just recognized “gawk” and ignored “reverse.” Somehow the definition becomes a look up label.
Begin with one clear line to set up who is who. “Two people appear. Max stands to the left, Mary to the right.” After that line, just use names.
Never link positions such as “left” or “right” in every line. When characters switch spots, every mention flips. Miss one and the entire story breaks. One setup line, fixed in one place, saves the day. Twenty random references guarantee trouble.
H3 appears to have character memory so if my right and left person cross – I can continue to prompt them as Max and Mary.
Two keys control the prompt’s audio. One key may steer the overall_soundscape. The other key operates non_diegetic_music – music that floats in but comes from no object in the scene.
You can switch off either key by writing N/A. Do not invent new terms. Only N/A works, straight from MiniMax’s own prompt notes. Ask for “None.” and H3 will probably guess. “No music please” does not work. “none” does not work.
Say nothing about music and H3 supplies a soundtrack. Every single time. Silence in your lines means nothing. If your aim is no music, non_diegetic_music: N/A becomes the only way. MiniMax’s own system prompt
You might call this a hunch, not hard law. Nine times out of ten, though, it stands up under testing.
If a sound links to an action you describe, mention it right where the action lives. The ambience key supplies the base world only. Imagine wind, water, people passing, cloth, or quiet breathing. To avoid surprises, attach each sound at its moment, then set both audio keys to N/A. That covers nearly all cases.
Leave the speaking lines blank and H3 does not remain silent. Instead, H3 makes up voice-like audio, no clear language. The result sounds something like parts of Romance tongues but never quite right. Impossible to use. Anyone who wants clear words must type those words.
Leave the lines unwritten and H3 does not go quiet. It invents speech-shaped audio in no identifiable language. It sounds vaguely Romance. It is unusable. If you want words, type the words.
Mouth movement seems to be the trickiest part of H3. The voice of one character might just leak to someone else, which has puzzled plenty of users. A known problem with the model sits at the heart of this struggle, not simply bad instructions. Discussions have popped up among developers searching for a reliable fix. open ComfyUI issue
H3 documentation says <d> marks speech. Inside the dialogue tag, place just the language marker and the words. Leave everything else outside. Speaker names, clues about how their voices sound, and even their delivery style must stand apart from the dialogue tag. Anything else might confuse the model.
Here is an example that probably helps: Max, older and deep-voiced, says: <d>[English] I told you it was cold.</d> Mary responds: <d>[English] Watch me.</d> The scene becomes clear. Voices connect to the right faces.
Now as a director, I’m highly resistant so what I use mostly is Max: “I told you it was cold.” Mary “No it is not!” and for me this seems to work ok – at least as often as the <d> framing. Right now I’m also placing the voice timbre queue in the first use of the voice “Max in a high squeeky voice: “I told you it was cold”.
H3 has a very annoying habit of having both people speak the same phrase at the same time.
Right now I’m testing this: Max: “I told you I was cold” Mary is silent. Mary “Watch me” – Max is silent.
Sometimes works. Sometimes in frustration I re-script a clip to match whoever H3 is favoring in the script. I don’t fight it, I join it.
When H3 has focused on one character, the other character may be partially on screen, or may be off screen. Assigning dialog to the off screen character does not work. So I prompt it separately. Voice from off screen “I mean it, it is very cold”. It seems to be based on whether the character’s lips are visible.
Write these words exactly—off-screen voice says:—and the line probably works perfectly. Use offscreen as one word and something goes wrong. The system matches full strings, not descriptions. That teaches a broader lesson. Phrases from H3’s own documentation should be copied precisely, character for character. Paraphrasing might break everything.
Early on, I made mistakes. Voice control seemed impossible until I figured out the placement. Where you put those details matters a lot. Many failed runs taught me: introduce details like age, pitch, and speaking speed at the first line someone delivers. Anywhere else, the model probably skips your instructions.
The model might guess the speaking language from a reference image, not the prompt. Pictures of people from East Asia, for example, somehow produced Mandarin even when I wrote lines in English. Language can slip away from your plan that quickly.
Anchor the language in your script. Adding a line such as “They speak English” usually keeps H3 on task. Hinting at a nationality sometimes works, but may also change the character’s look—so that trick remains unpredictable.
H3 will probably not slow down to fill the moments you want. Give H3 more seconds than action beats and it might just invent extra scenes or repeat content. The best results come from matching your video’s timing to the number of actions you planned.
Physics looks believable for single actions. A basic throw always works. Try combining a throw with several bounces and H3 probably fails, in every phrasing I tried. Breaking big actions into smaller moments gives much better results.
Sometimes, actions run in reverse. Once, the ball traveled backwards into a hand instead of away. The model might have learned this as normal from reversed footage in training, so no default direction exists. Describing direction clearly—“away from the body, forward into the distance”—locks the motion in. Later references keep the action moving as planned.
H3 takes hints from what you feed it. I have seen raised arms in the first frame become a throw before prompts even begin. Picking a frame where everyone looks relaxed and neutral may prevent these surprise movements. The starting moment shapes everything that follows.
H3 reads intent out of your source image. A raised arm in my still became a throw, before the prompted action even started. Choose start frames with neutral posture unless you want that motion.
One time, I tried to use [Shot 2] and [Shot 3] for splitting story moments. That was a mistake. Those tools tell the system when to cut scenes. Unwanted scene breaks kept popping up when I just wanted smooth flow. Now, every action lives together in a single Shot 1 for me.
Many hours disappeared from my life because of this mixup. Sometimes a video clip only shows what you want for a few heartbeats, then everything changes randomly. Faces might lose character. Backgrounds probably shift and look new.
No magic writing trick solves that puzzle. For days, I experimented. The real problem lies in the step limits, and your chosen turbo model changes the result. Workflow challenges make more trouble than writing mistakes. That topic probably needs a whole post for itself.
When writing, go long and use simple language. Describe every action just once, then call it by its chosen name later. Put all names and their spots right at the beginning. Set both audio options to N/A, unless you really want some surprises. Use quotes for each line people speak. Always choose the d tag. At a character’s first words, mention the voice. Clearly say the language. Keep one movement or action for each short clip. Try to match every clip’s length to your dramatic beats. Throw out the shot markers.
None of these secret moves appear in the official instructions. Bad drafts taught me each one. Anyone planning to share work made with H3 should check the terms carefully. People in the United States and the European Union might find a real problem hidden in the rules.
A year ago, I blogged that the rate of Video AI sophistication is moving so…
Property managers who prioritize systematic care safeguard enterprise balance sheets and maintain daily operational stability.…
Buying headphones online or in store? This guide covers what to check, whether you end…
Understanding how charging power, battery level, temperature, and driving patterns interact can help drivers choose…
Building an impactful marketing approach requires continuous refinement and a commitment to understanding your audience.…
You still need shade, a stable surface, reliable Wi-Fi, and somewhere protected enough for your…