Supported Models

Explore the video, image, and text AI models supported in CawCut.

Updated Aug 20, 2026 · 6 min read

CawCut includes a growing library of image, video, and text models. You can use them for generation, editing, prompt development, workflow logic, and other creative tasks.

You can combine models from different providers in one workflow or replace a model node as your needs change.

1. Video Models

To generate clips or connect video inputs and outputs in a workflow, see Video Generation.

AI ModelCore Strengths & Use Cases
Seedance 2.5ByteDance’s next-generation audio-video model for coherent stories up to 30 seconds, with precise multimodal reference control and advanced editing.
Best for: Longer narrative videos, reference-driven creation, and professional video production.
Seedance 2.0ByteDance’s multi-modal powerhouse supporting 4K output. Features advanced capabilities like multi-modal reference inputs (text, image, audio) and director-level camera control for cinematic motion.
Best for: Cinematic scenes, advertising videos, and professional video production.
Seedance 2.0 FastA faster Seedance 2.0 option for shorter turnaround while retaining strong motion and visual quality.
Best for: Rapid iteration, previews, and social content.
Seedance 2.0 MiniA lightweight Seedance option for efficient video generation and workflow testing.
Best for: Quick drafts, simple motion, and early-stage concepts.
Kling 3.0 TurboA speed-optimized Kling 3.0 model that balances strong motion, prompt adherence, and efficient generation.
Best for: Fast iterations, social videos, and concept development.
Kling 3.0 Turbo ProThe higher-quality Turbo option for greater detail, consistency, and complex motion while retaining faster generation.
Best for: Polished ads, cinematic sequences, and production-ready content.
Kling 3.0 Turbo StandardAn efficient Turbo option for stable everyday text-to-video and image-to-video generation.
Best for: Routine content creation, previews, and batch iteration.
Kling 3.0 ProA high-performance model engineered for cinematic storytelling. Excels in character animation, precise motion transfer, and seamless multi-shot generation.
Best for: Commercials, short films, music videos, and branded content.
Kling 3.0 StandardDelivers highly stable text-to-video and image-to-video generation with robust audio synchronization.
Best for: General creative tasks and daily video content production.
Google Veo 3.1Google’s flagship video model, renowned for highly realistic physics, authentic visual textures, and profound cinematic depth.
Best for: High-end commercial visuals and conceptual storytelling.
Google Veo 3.1 FastAn accelerated version of Veo, slashing generation times while maintaining exceptional core visual quality.
Best for: Rapid content iteration, batch processing, and social media shorts.
Hailuo 2.3Hailuo AI’s video model for generating expressive movement and polished short-form visuals.
Best for: Character motion, creative clips, and social video concepts.

2. Image Models

To generate, edit, or pass images between nodes, see Image Generation Nodes.

AI ModelCore Strengths & Use Cases
GPT Image 2OpenAI’s powerhouse model, featuring exceptional prompt understanding and visual fidelity.
Best for: Creative visuals, poster design, and concept art.
GPT Image 1A versatile core model for reliable text-to-image generation and granular image editing.
Best for: Daily design tasks, illustrations, and marketing materials.
GPT Image 1 miniA lightweight OpenAI image model for faster, lower-cost visual generation and iteration.
Best for: Quick concepts, simple assets, and workflow testing.
Nano Banana 2 (Gemini 3.1 Flash)A lightning-fast model deeply optimized for lightweight tasks and rapid generation.
Best for: Social media assets and high-speed visual ideation.
Nano Banana Pro (Gemini 3 Pro)Flagship performance offering profound semantic understanding and intricate detail rendering.
Best for: Complex scene construction and commercial-grade visual output.
Nano BananaA versatile Google image model for prompt-based generation and image editing.
Best for: Everyday creative tasks, visual exploration, and image variations.
Flux 2 ProBlack Forest Labs’ professional image model for detailed, polished visual generation.
Best for: High-quality campaign visuals, concept art, and commercial assets.
Seedream 5.0 ProByteDance’s advanced Seedream model for high-quality image generation and demanding creative work.
Best for: Detailed compositions, production visuals, and commercial design.
Seedream 4.5ByteDance’s premium model, perfectly aligned with Asian aesthetics and deep Chinese context understanding.
Best for: E-commerce visuals, portrait photography, and posters.
Seedream 5.0 liteA lightweight, next-generation model delivering improvements in character consistency and spatial layout reasoning.
Best for: Faster iteration, e-commerce graphics, and layout exploration.
Ideogram 4.0Ideogram’s image model for design-oriented generation and text-rich compositions.
Best for: Posters, branded graphics, typography, and advertising concepts.
Reve ImageReve’s image generation model for creating polished visuals from natural-language prompts.
Best for: Creative ideation, illustrations, and marketing imagery.
Z Image TurboTongyi’s fast image model for rapid visual generation and experimentation.
Best for: Quick drafts, batch concepts, and fast iteration.
Recraft V3Recraft’s design-focused model for consistent, editable-looking visual assets.
Best for: Brand graphics, illustrations, icons, and design systems.
Wan 2.6Tongyi’s visual engine, boasting strong artistic styling capabilities and precise language interpretation.
Best for: Concept design and stylized artistic illustrations.

3. Text Models

Text models can help develop prompts, structure ideas, analyze inputs, and support text-based steps inside a workflow.

To add and configure text-based workflow steps, see Text Generation Nodes.

ProviderAvailable Models
OpenAIGPT 5.5, GPT 5.4, GPT 5.4 mini, GPT 5.4 nano, GPT 5.2, GPT 5.1, GPT 5 mini, GPT 5 nano
AnthropicOpus 4.8, Opus 4.7, Opus 4.6, Sonnet 4.6, Haiku 4.5
DeepSeekDeepSeek V4 Pro, DeepSeek V4 Flash

Related articles

Frequently Asked Questions

Use these tips to choose the right model for your workflow.

Use a flagship or Pro model when detail, consistency, or complex instructions matter most. Use a Fast, Turbo, Mini, lite, nano, Flash, or other lightweight option for drafts and rapid iteration, then switch to a higher-quality model for the final result if needed.

Start with Seedance for narrative or reference-driven work, Kling for character motion and polished creative video, Veo for realism and cinematic depth, or Hailuo for expressive short-form clips. Choose a Fast, Turbo, Standard, or Mini variant when turnaround matters more than maximum quality.

Try Ideogram for text-rich graphics, Recraft for brand assets and design systems, and Nano Banana 2, Z Image Turbo, or GPT Image 1 mini for quick exploration. For detailed commercial visuals, start with GPT Image 2, Nano Banana Pro, Flux 2 Pro, or Seedream 5.0 Pro.

Use larger or Pro models for complex planning, analysis, and detailed prompt development. Use mini, nano, Haiku, or Flash models for simple transformations, high-volume steps, and faster iteration. Test important prompts with a few models because writing style and instruction following can differ.

Yes. You can connect compatible image, video, text, and processing nodes from different providers in one workflow, then replace or adjust individual nodes as your project changes.