Master MiniMax H3 Emotion & Speech Tags in CawCut

Learn how to use inline speech tags, emotional context, and microexpressions to create more expressive MiniMax H3 videos in CawCut.

Updated Sep 10, 2026 · 5 min read

MiniMax H3 can generate more expressive performances when the prompt describes not only what a character says, but also how the line is delivered and what the character does while speaking. In CawCut, you can combine inline speech cues with clear visual behavior to guide laughter, breaths, emphasis, hesitation, surprise, and other reactions. Results can vary, so test short prompts first.

1. Understand How Tags Work

Use four layers of direction:

  • Speech formatting tags: Inline tags placed inside spoken text, such as <pause>, <inhale>, <whisper>...</whisper>, and <i>word</i>.
  • Vocal sound-effect tags: Preset cues written with angle brackets, such as <chuckle>, <sighs>, or <gasps>, which can add non-verbal sounds to the generated audio.
  • Emotional context: Natural language outside the dialogue that explains the character’s intention, mood, or delivery.
  • Visible behavior: Facial microexpressions and body actions that make the emotion visible, such as a brief smile, tightened jaw, raised eyebrows, or a pause before eye contact.

Tags are best treated as local timing cues. Place each cue close to the word or moment it should affect. Do not rely on a tag alone to create a complete emotion.

For example, this is too vague:

The woman feels nervous.

This gives MiniMax H3 more visible information:

The woman checks her watch twice, taps her fingers against the folder, takes a short breath, and looks toward the closed door.

2. Write the Dialogue Block

MiniMax H3 uses a <d> block for spoken content. Keep the language label and the exact spoken words inside the block. Put the speaker’s identity, action, and delivery description outside it.

Use a stable speaker ID whenever a character speaks. Keep the same ID for that character in every shot.

The young woman with a quiet, breathy voice (S1) says hesitantly:
<d>[English] I thought you would be here. <inhale> I was wrong.</d>

You can add an emphasis cue near a specific word:

The man (S1) leans toward the camera and speaks with controlled anger:
<d>[English] You promised me this would never happen. <i>Never.</i></d>

You can also add a non-verbal reaction:

The woman (S1) tries to stay serious, then breaks into a small laugh:
<d>[English] I cannot believe you did that. <chuckle></d>

Speech formatting cues include:

CueIntended useExample
<pause>Short pauseThe signal is fading. <pause> Stay with me.
<long pause>Longer pauseShe asked me to leave... <long pause> and I stayed.
<breath>Breathing soundHe reaches the top floor. <breath> The door is open.
<inhale> / <exhale>Inward or outward breath<inhale> We begin on three.
<catches breath>Out-of-breath deliveryI ran all the way here. <catches breath> Give me a moment.
<deep breath>Deep breath used to calm down<deep breath> I am ready to face them.
<i>word</i>Emphasize one to four wordsThis is <i>exactly</i> what I warned you about.
<whisper>...</whisper>Whisper the enclosed words<whisper> The lights are still on.</whisper>
<stutter>Stutter or broken delivery<stutter> N-no, that was not my idea.

These vocal sound-effect cues are preset-style experiments. Custom emotion labels written with angle brackets, such as <nervous>, may be read as literal speech rather than interpreted as controls.

CueWhat it doesExample
<laughs> / <chuckle>Adds laughter, from a stronger laugh to a brief amused reactionThat is the funniest thing I have heard. <laughs>
<snorts>Adds a snort, often as an involuntary amused reactionI tried not to laugh. <snorts>
<humming>Adds humming instead of spoken wordsShe walks down the hallway <humming>
<breath>Adds a general breathing soundAnd then... <breath> it happened.
<inhale> / <exhale>Adds an inward or outward breath<inhale> All right, let us do this.
<pant> / <gasps>Adds rapid panting or a sharp gaspHe reaches the finish line <gasps> at last.
<coughs> / <clear-throat>Adds coughing or throat-clearing before speech<clear-throat> I have an announcement.
<sighs>Adds a sigh that can suggest resignation, sadness, or reliefI suppose you are right. <sighs>
<lip-smacking>Adds a lip-smacking mouth soundHe pauses, <lip-smacking> then speaks.
<sniffs>Adds sniffingShe wipes her eyes and <sniffs>.
<groans>Adds a groan, such as discomfort or frustration<groans> I cannot lift this.
<yawns>Adds a yawnIt is very late. <yawns>
<sneezes>Adds a sneezeHe looks toward the flowers and <sneezes>.
<burps>Adds a burpThe character finishes the drink and <burps>.
<whistles>Adds whistlingHe walks away <whistles>
<hissing>Adds a hissing soundThe steam escapes with a sharp <hissing>.
<emm>Adds an “emm” hesitation or filler sound<emm> I am not sure.
<crying>Adds crying or sobbingI tried to stop it. <crying>
<applause>Adds applause in the scene or backgroundThe speaker finishes, followed by <applause>.

These cues work best as precise, local instructions rather than as a separate vocabulary to memorize. Use the complete speaker format—(S1) plus [English] inside <d>—and keep each cue close to the moment it should affect. For emphasis, use <i>word</i> around one to four words; square-bracket cues such as [emphasis] can sometimes be spoken aloud or apply to too much of the sentence. If a cue produces an unwanted result, replace it with a natural-language delivery direction outside the dialogue block. When several emphasis cues are needed, separate them with <pause> so the model has room to distinguish each beat.

3. Add Visible Emotion and Context

Speech tags can influence audio, but the face and body still need visual direction. Describe the reaction in the shot timeline:

[Shot 1] Cinematic close-up of a tired detective in a dim office. He reads the message, raises his eyebrows for a moment, then presses his lips together and looks away. His voice is restrained but hurt (S1):
<d>[English] You promised to stay. <sighs> What am I supposed to do now?</d>

Try this scene yourself!

Use concrete, observable details:

  • Surprise: eyes widen, eyebrows lift, mouth opens briefly, head pulls back.
  • Doubt: one eyebrow rises, gaze shifts sideways, lips tighten.
  • Fear: shallow breathing, widened eyes, tense shoulders, a quick glance toward the exit.
  • Relief: shoulders drop, eyelids soften, a long exhale, a small uneven smile.
  • Anger: tightened jaw, narrowed eyes, rigid posture, clipped delivery.
  • Sadness: lowered gaze, slower movement, trembling lips, a quiet breath before speaking.

Context makes the emotion more consistent. Explain what happened immediately before the line and what the character wants. “She is afraid because the door handle is moving” usually guides the performance better than “she is afraid.”

4. Build and Test the Prompt

  1. Open the MiniMax H3 video generation workflow in CawCut.
  2. Choose the input mode you need, such as text-to-video or image-to-video.
  3. Write the shot description in chronological order: setting, subject, action, camera, reaction, and dialogue.
  4. Add a stable speaker ID and put the exact spoken line inside <d>.
  5. Add one or two inline cues at the moment they should occur. Use <i>...</i> for one to four emphasized words, formatting tags for timing or delivery, and preset angle-bracket cues for vocal sounds.
  6. Generate a short, lower-cost test when available.
  7. Check the voice, timing, facial reaction, mouth movement, and continuity before expanding the prompt or increasing output quality.

Keep the prompt focused. One emotional change per shot is usually easier to control than a character who moves from fear to anger, laughter, and sadness in the same sentence.

5. Troubleshoot Unreliable Results

  • The tag is spoken aloud: Use the complete (S1) says: <d>[English] ...</d> structure, try a different syntax, or replace the tag with natural-language direction outside <d>. Isolated tags can sometimes be verbalized.
  • The emotion sounds right but looks wrong: Add specific facial and body actions to the shot description.
  • The character skips the breath or laugh: Move the cue closer to the affected phrase, use the matching preset such as <breath> or <laughs>, and reduce the number of other tags.
  • The character changes voice between shots: Reuse the same speaker ID and voice description, and keep the dialogue structure consistent.
  • The line is rushed: Shorten the sentence, extend the generation time, add a pause or breath, and reduce the number of actions in the shot.
  • The result is inconsistent: Change one variable at a time. Test the same shot with no tag, one tag, and then a different natural-language description.

For especially expressive faces, test a close-up with a simple background and keep the subject’s head and shoulders stable in frame. This gives the model fewer competing visual tasks while it generates speech and microexpressions.

The most reliable prompt is a small production brief: describe what the viewer sees, what the character does, what the character says, and how the sound changes at that moment.

Related articles

Frequently asked questions

Find answers about using emotion and delivery cues with MiniMax H3.

Put the cue directly in the spoken dialogue inside the <d> block, immediately before or around the words it should affect. Keep the speaker identity, action, and delivery description outside the block. For emphasis, use <i>word</i> around one to four words.

Use the complete speaker format, such as (S1) says: <d>[English] ...</d>, instead of testing an isolated tag. For emphasis, use <i>word</i> around one to four words rather than [emphasis]. If the result is still unreliable, describe the delivery naturally outside the dialogue block.

Speech cues primarily influence delivery and non-verbal sound. Describe the visible reaction separately—for example, narrowed eyes, a hesitant smile, a lowered gaze, or a slow exhale—and place it in the shot description.

Yes, but start with one or two cues at meaningful points. Leave space between closely stacked cues, and add a pause when needed. Too many tags can compete with the dialogue and make timing less predictable.

Give the speaking character a stable speaker ID such as (S1), repeat the same identity and voice description, and describe each emotional change in chronological order.