MiniMax H3 can generate more expressive performances when the prompt describes not only what a character says, but also how the line is delivered and what the character does while speaking. In CawCut, you can combine inline speech cues with clear visual behavior to guide laughter, breaths, emphasis, hesitation, surprise, and other reactions. Results can vary, so test short prompts first.
1. Understand How Tags Work
Use four layers of direction:
- Speech formatting tags: Inline tags placed inside spoken text, such as
<pause>,<inhale>,<whisper>...</whisper>, and<i>word</i>. - Vocal sound-effect tags: Preset cues written with angle brackets, such as
<chuckle>,<sighs>, or<gasps>, which can add non-verbal sounds to the generated audio. - Emotional context: Natural language outside the dialogue that explains the character’s intention, mood, or delivery.
- Visible behavior: Facial microexpressions and body actions that make the emotion visible, such as a brief smile, tightened jaw, raised eyebrows, or a pause before eye contact.
Tags are best treated as local timing cues. Place each cue close to the word or moment it should affect. Do not rely on a tag alone to create a complete emotion.
For example, this is too vague:
The woman feels nervous.This gives MiniMax H3 more visible information:
The woman checks her watch twice, taps her fingers against the folder, takes a short breath, and looks toward the closed door.2. Write the Dialogue Block
MiniMax H3 uses a <d> block for spoken content. Keep the language label and the exact spoken words inside the block. Put the speaker’s identity, action, and delivery description outside it.
Use a stable speaker ID whenever a character speaks. Keep the same ID for that character in every shot.
The young woman with a quiet, breathy voice (S1) says hesitantly:
<d>[English] I thought you would be here. <inhale> I was wrong.</d>You can add an emphasis cue near a specific word:
The man (S1) leans toward the camera and speaks with controlled anger:
<d>[English] You promised me this would never happen. <i>Never.</i></d>You can also add a non-verbal reaction:
The woman (S1) tries to stay serious, then breaks into a small laugh:
<d>[English] I cannot believe you did that. <chuckle></d>Speech formatting cues include:
| Cue | Intended use | Example |
|---|---|---|
<pause> | Short pause | The signal is fading. <pause> Stay with me. |
<long pause> | Longer pause | She asked me to leave... <long pause> and I stayed. |
<breath> | Breathing sound | He reaches the top floor. <breath> The door is open. |
<inhale> / <exhale> | Inward or outward breath | <inhale> We begin on three. |
<catches breath> | Out-of-breath delivery | I ran all the way here. <catches breath> Give me a moment. |
<deep breath> | Deep breath used to calm down | <deep breath> I am ready to face them. |
<i>word</i> | Emphasize one to four words | This is <i>exactly</i> what I warned you about. |
<whisper>...</whisper> | Whisper the enclosed words | <whisper> The lights are still on.</whisper> |
<stutter> | Stutter or broken delivery | <stutter> N-no, that was not my idea. |
These vocal sound-effect cues are preset-style experiments. Custom emotion labels written with angle brackets, such as <nervous>, may be read as literal speech rather than interpreted as controls.
| Cue | What it does | Example |
|---|---|---|
<laughs> / <chuckle> | Adds laughter, from a stronger laugh to a brief amused reaction | That is the funniest thing I have heard. <laughs> |
<snorts> | Adds a snort, often as an involuntary amused reaction | I tried not to laugh. <snorts> |
<humming> | Adds humming instead of spoken words | She walks down the hallway <humming> |
<breath> | Adds a general breathing sound | And then... <breath> it happened. |
<inhale> / <exhale> | Adds an inward or outward breath | <inhale> All right, let us do this. |
<pant> / <gasps> | Adds rapid panting or a sharp gasp | He reaches the finish line <gasps> at last. |
<coughs> / <clear-throat> | Adds coughing or throat-clearing before speech | <clear-throat> I have an announcement. |
<sighs> | Adds a sigh that can suggest resignation, sadness, or relief | I suppose you are right. <sighs> |
<lip-smacking> | Adds a lip-smacking mouth sound | He pauses, <lip-smacking> then speaks. |
<sniffs> | Adds sniffing | She wipes her eyes and <sniffs>. |
<groans> | Adds a groan, such as discomfort or frustration | <groans> I cannot lift this. |
<yawns> | Adds a yawn | It is very late. <yawns> |
<sneezes> | Adds a sneeze | He looks toward the flowers and <sneezes>. |
<burps> | Adds a burp | The character finishes the drink and <burps>. |
<whistles> | Adds whistling | He walks away <whistles> |
<hissing> | Adds a hissing sound | The steam escapes with a sharp <hissing>. |
<emm> | Adds an “emm” hesitation or filler sound | <emm> I am not sure. |
<crying> | Adds crying or sobbing | I tried to stop it. <crying> |
<applause> | Adds applause in the scene or background | The speaker finishes, followed by <applause>. |
These cues work best as precise, local instructions rather than as a separate vocabulary to memorize. Use the complete speaker format—(S1) plus [English] inside <d>—and keep each cue close to the moment it should affect. For emphasis, use <i>word</i> around one to four words; square-bracket cues such as [emphasis] can sometimes be spoken aloud or apply to too much of the sentence. If a cue produces an unwanted result, replace it with a natural-language delivery direction outside the dialogue block. When several emphasis cues are needed, separate them with <pause> so the model has room to distinguish each beat.
3. Add Visible Emotion and Context
Speech tags can influence audio, but the face and body still need visual direction. Describe the reaction in the shot timeline:
[Shot 1] Cinematic close-up of a tired detective in a dim office. He reads the message, raises his eyebrows for a moment, then presses his lips together and looks away. His voice is restrained but hurt (S1):
<d>[English] You promised to stay. <sighs> What am I supposed to do now?</d>Use concrete, observable details:
- Surprise: eyes widen, eyebrows lift, mouth opens briefly, head pulls back.
- Doubt: one eyebrow rises, gaze shifts sideways, lips tighten.
- Fear: shallow breathing, widened eyes, tense shoulders, a quick glance toward the exit.
- Relief: shoulders drop, eyelids soften, a long exhale, a small uneven smile.
- Anger: tightened jaw, narrowed eyes, rigid posture, clipped delivery.
- Sadness: lowered gaze, slower movement, trembling lips, a quiet breath before speaking.
Context makes the emotion more consistent. Explain what happened immediately before the line and what the character wants. “She is afraid because the door handle is moving” usually guides the performance better than “she is afraid.”
4. Build and Test the Prompt
- Open the MiniMax H3 video generation workflow in CawCut.
- Choose the input mode you need, such as text-to-video or image-to-video.
- Write the shot description in chronological order: setting, subject, action, camera, reaction, and dialogue.
- Add a stable speaker ID and put the exact spoken line inside
<d>. - Add one or two inline cues at the moment they should occur. Use
<i>...</i>for one to four emphasized words, formatting tags for timing or delivery, and preset angle-bracket cues for vocal sounds. - Generate a short, lower-cost test when available.
- Check the voice, timing, facial reaction, mouth movement, and continuity before expanding the prompt or increasing output quality.
Keep the prompt focused. One emotional change per shot is usually easier to control than a character who moves from fear to anger, laughter, and sadness in the same sentence.
5. Troubleshoot Unreliable Results
- The tag is spoken aloud: Use the complete
(S1) says: <d>[English] ...</d>structure, try a different syntax, or replace the tag with natural-language direction outside<d>. Isolated tags can sometimes be verbalized. - The emotion sounds right but looks wrong: Add specific facial and body actions to the shot description.
- The character skips the breath or laugh: Move the cue closer to the affected phrase, use the matching preset such as
<breath>or<laughs>, and reduce the number of other tags. - The character changes voice between shots: Reuse the same speaker ID and voice description, and keep the dialogue structure consistent.
- The line is rushed: Shorten the sentence, extend the generation time, add a pause or breath, and reduce the number of actions in the shot.
- The result is inconsistent: Change one variable at a time. Test the same shot with no tag, one tag, and then a different natural-language description.
For especially expressive faces, test a close-up with a simple background and keep the subject’s head and shoulders stable in frame. This gives the model fewer competing visual tasks while it generates speech and microexpressions.
The most reliable prompt is a small production brief: describe what the viewer sees, what the character does, what the character says, and how the sound changes at that moment.