Video Generation

Learn how to use the Video Generation node to create videos from text, images, frames, audio, and video references.

2026년 8월 31일 업데이트 · 8분 읽기

The Video Generation node creates new video clips from text prompts and media references. Choose a generation mode based on the inputs you have, select a compatible model, configure the output, and click [Run]. The completed video is passed to the Video output port so it can be connected to other nodes in the workflow.

1. Generation Modes

Select a mode from the tabs at the top of the Node Panel. The node’s input ports and quick-add controls change to match the selected mode.

1.1 Text to Video

Use [Text to Video] when you want to create a complete scene from a written description without supplying a visual reference.

Video Generation node and Node Panel with Text to Video selected.
  • Prompt input: Connect a Text node to the Prompt port, or enter the description directly in the prompt box.
  • Best for: Exploring new concepts, establishing shots, stylized scenes, and clips whose composition can be described entirely with text.

1.2 Image to Video

Use [Image to Video] to animate a still image or guide the generated clip with visual references.

Video Generation node and Node Panel with Image to Video selected.
  • Prompt input: Describe the subject’s movement, camera movement, scene changes, and any details that should remain consistent.
  • Image input: Connect one or more images through the Image port, or use [Image] in the Node Panel to add a reference. The maximum number depends on the selected model.
  • Best for: Animating artwork or product images, preserving a subject’s appearance, and adding controlled motion to an existing composition.

1.3 First/Last Frame to Video

Use [First/Last Frame to Video] when the clip must begin from a specific image and optionally finish on another image.

Video Generation node and Node Panel with First/Last Frame to Video selected.
  • **First Frame\*:** Connect the required starting image to define the opening composition.
  • Last Frame: Connect an optional ending image to guide where the movement, camera, or transition should finish.
  • Prompt input: Describe how the subject and camera should move between the supplied frames.
  • Best for: Planned transitions, before-and-after shots, controlled camera moves, and sequences that must arrive at a defined final composition.

1.4 Omni to Video

Use [Omni to Video] to combine several media types in one multimodal request.

Video Generation node and Node Panel with Omni to Video selected.
  • Prompt input: Explain how the connected references should influence the generated result.
  • Image inputs: Connect images as subject, character, style, or composition references.
  • Video inputs: Connect videos as motion, performance, or scene references.
  • Audio inputs: Connect audio files to guide timing, rhythm, speech, or sound.
  • Quick-add controls: Use [Image], [Video], and [Audio] in the Node Panel to add references directly.
  • Best for: Reference-rich scenes that need coordinated visuals, movement, timing, and audio.

2. Node Settings

The controls at the bottom of the Node Panel determine how the video is generated. Available values can change when you switch models or generation modes because each model supports different inputs and output options.

2.1 Choose a Model

Select the model menu to choose the generation engine. Use the table below as a starting point; Fast, Mini, Standard, and Turbo tiers are useful for iteration, while full and Pro tiers generally prioritize final quality and control.

ModelBest for
Seedance 2.530-second audiovisual storytelling that needs precise multimodal reference control and targeted video editing.
Seedance 2.0Multimodal, cinematic scenes that combine strong motion, reference media, and synchronized audio.
Seedance 2.0 FastFaster multimodal iterations when you want to test a scene without leaving the Seedance workflow.
Seedance 2.0 MiniQuick drafts and lower-cost prompt tests before committing to a final generation.
Kling 3.0 ProHigh-detail cinematic clips, longer or more complex action, strong prompt adherence, and polished audio.
Kling 3.0 StandardGeneral-purpose video generation with a balance of quality, speed, and credit use.
Kling 3.0 Turbo ProFaster generations that still prioritize detail and production-ready visual quality.
Kling 3.0 Turbo StandardRapid, economical drafts for testing motion, framing, and prompt direction.
Veo 3.1Cinematic realism, strong prompt adherence, and polished audiovisual scenes.
Veo 3.1 FastQuicker Veo iterations when turnaround matters more than maximum output quality.

2.2 Model Specifications

ModelDurationAspect ratioResolutionGenerate audioPrompt lengthMaximum reference inputs
Seedance 2.54–30 secondsAuto, 16:9, 9:16, 1:1, 21:9, 4:3, 3:4480P, 720P3–20,000 characters30 images; 10 videos; 10 audio files
Seedance 2.04–15 secondsAuto, 16:9, 9:16, 1:1, 21:9, 4:3, 3:4480P, 720P, 1080P, 4KUp to 20,000 characters9 images; 3 videos; 3 audio files
Seedance 2.0 Fast/Mini4–15 secondsAuto, 16:9, 9:16, 1:1, 21:9, 4:3, 3:4480P, 720P3–20,000 characters9 images; 3 videos; 3 audio files
Kling 3.0 Pro/Standard3–15 seconds16:9, 9:16, 1:1Not selectable3–2,500 characters4 images; 2 first/last frames
Kling 3.0 Turbo Pro/Standard3–15 seconds16:9, 9:16, 1:1Not selectableUp to 3,072 characters1 image
Veo 3.1/3.1 Fast4, 6, or 8 seconds16:9, 9:16720P, 1080PUp to 10,000 characters1 image

2.3 Model Support by Generation Mode

ModelText to VideoImage to VideoFirst/Last FrameOmni to Video
Seedance 2.0/2.5
Kling 3.0 Pro/Standard
Kling 3.0 Turbo Pro/Standard
Veo 3.1/3.1 Fast

2.4 Output Settings

Configure the output controls after selecting a model. If a value is unavailable, that model does not support it in the current generation mode.

SettingWhat it doesHow to use it
DurationSets the length of the generated clip.Drag the slider or enter an available duration. Longer videos generally require more generation time and credits.
RatioSets the video’s frame shape.Choose [Auto], [16:9], [9:16], [1:1], [21:9], [4:3], or [3:4] when available. [Auto] lets the model follow the reference or its default framing.
ResolutionSets the output dimensions and level of visible detail.Choose [480P], [720P], [1080P], or [4K] when supported. Higher resolutions can take longer and use more credits.
Generate AudioControls whether the model creates an accompanying audio track.Switch it [On] for generated sound or [Off] for a silent clip. Audio generation is only available on supported models.
Advanced Settings > SeedSets the random starting value used for the generation.Keep the same prompt, reference inputs, model, output settings, and seed for a more similar variation. Enter a different value or use the refresh control to explore a new result.
Credit estimateShows the estimated credits for the next generation.Review the number beside the credit icon. It updates when you change the model or output settings.
RunStarts the video generation with the current inputs and settings.Check the prompt, references, and credit estimate, then click [Run].

관련 기사

Frequently asked questions

Find answers about Video Generation modes, models, and output settings.

Each model supports a different combination of inputs. Select the generation mode first, then choose from the compatible models shown in the model menu.

Use a Fast, Mini, Standard, or Turbo option to test the prompt and motion more quickly. After the composition works, switch to a higher-quality model or tier for the final generation.

The available settings depend on the selected model and generation mode. CawCut only displays or enables combinations supported by that model.

The seed provides the model with a repeatable random starting point. Reusing the same prompt, model, settings, and seed can produce a more similar result, while changing the seed helps create a new variation.