CawCut includes a growing library of image, video, and text models. You can use them for generation, editing, prompt development, workflow logic, and other creative tasks.
You can combine models from different providers in one workflow or replace a model node as your needs change.
1. Video Models
To generate clips or connect video inputs and outputs in a workflow, see Video Generation.
| AI Model | Core Strengths & Use Cases |
|---|---|
| Seedance 2.5 | ByteDance’s next-generation audio-video model for coherent stories up to 30 seconds, with precise multimodal reference control and advanced editing. Best for: Longer narrative videos, reference-driven creation, and professional video production. |
| Seedance 2.0 | ByteDance’s multi-modal powerhouse supporting 4K output. Features advanced capabilities like multi-modal reference inputs (text, image, audio) and director-level camera control for cinematic motion. Best for: Cinematic scenes, advertising videos, and professional video production. |
| Seedance 2.0 Fast | A faster Seedance 2.0 option for shorter turnaround while retaining strong motion and visual quality. Best for: Rapid iteration, previews, and social content. |
| Seedance 2.0 Mini | A lightweight Seedance option for efficient video generation and workflow testing. Best for: Quick drafts, simple motion, and early-stage concepts. |
| Kling 3.0 Turbo | A speed-optimized Kling 3.0 model that balances strong motion, prompt adherence, and efficient generation. Best for: Fast iterations, social videos, and concept development. |
| Kling 3.0 Turbo Pro | The higher-quality Turbo option for greater detail, consistency, and complex motion while retaining faster generation. Best for: Polished ads, cinematic sequences, and production-ready content. |
| Kling 3.0 Turbo Standard | An efficient Turbo option for stable everyday text-to-video and image-to-video generation. Best for: Routine content creation, previews, and batch iteration. |
| Kling 3.0 Pro | A high-performance model engineered for cinematic storytelling. Excels in character animation, precise motion transfer, and seamless multi-shot generation. Best for: Commercials, short films, music videos, and branded content. |
| Kling 3.0 Standard | Delivers highly stable text-to-video and image-to-video generation with robust audio synchronization. Best for: General creative tasks and daily video content production. |
| Google Veo 3.1 | Google’s flagship video model, renowned for highly realistic physics, authentic visual textures, and profound cinematic depth. Best for: High-end commercial visuals and conceptual storytelling. |
| Google Veo 3.1 Fast | An accelerated version of Veo, slashing generation times while maintaining exceptional core visual quality. Best for: Rapid content iteration, batch processing, and social media shorts. |
| Hailuo 2.3 | Hailuo AI’s video model for generating expressive movement and polished short-form visuals. Best for: Character motion, creative clips, and social video concepts. |
2. Image Models
To generate, edit, or pass images between nodes, see Image Generation Nodes.
| AI Model | Core Strengths & Use Cases |
|---|---|
| GPT Image 2 | OpenAI’s powerhouse model, featuring exceptional prompt understanding and visual fidelity. Best for: Creative visuals, poster design, and concept art. |
| GPT Image 1 | A versatile core model for reliable text-to-image generation and granular image editing. Best for: Daily design tasks, illustrations, and marketing materials. |
| GPT Image 1 mini | A lightweight OpenAI image model for faster, lower-cost visual generation and iteration. Best for: Quick concepts, simple assets, and workflow testing. |
| Nano Banana 2 (Gemini 3.1 Flash) | A lightning-fast model deeply optimized for lightweight tasks and rapid generation. Best for: Social media assets and high-speed visual ideation. |
| Nano Banana Pro (Gemini 3 Pro) | Flagship performance offering profound semantic understanding and intricate detail rendering. Best for: Complex scene construction and commercial-grade visual output. |
| Nano Banana | A versatile Google image model for prompt-based generation and image editing. Best for: Everyday creative tasks, visual exploration, and image variations. |
| Flux 2 Pro | Black Forest Labs’ professional image model for detailed, polished visual generation. Best for: High-quality campaign visuals, concept art, and commercial assets. |
| Seedream 5.0 Pro | ByteDance’s advanced Seedream model for high-quality image generation and demanding creative work. Best for: Detailed compositions, production visuals, and commercial design. |
| Seedream 4.5 | ByteDance’s premium model, perfectly aligned with Asian aesthetics and deep Chinese context understanding. Best for: E-commerce visuals, portrait photography, and posters. |
| Seedream 5.0 lite | A lightweight, next-generation model delivering improvements in character consistency and spatial layout reasoning. Best for: Faster iteration, e-commerce graphics, and layout exploration. |
| Ideogram 4.0 | Ideogram’s image model for design-oriented generation and text-rich compositions. Best for: Posters, branded graphics, typography, and advertising concepts. |
| Reve Image | Reve’s image generation model for creating polished visuals from natural-language prompts. Best for: Creative ideation, illustrations, and marketing imagery. |
| Z Image Turbo | Tongyi’s fast image model for rapid visual generation and experimentation. Best for: Quick drafts, batch concepts, and fast iteration. |
| Recraft V3 | Recraft’s design-focused model for consistent, editable-looking visual assets. Best for: Brand graphics, illustrations, icons, and design systems. |
| Wan 2.6 | Tongyi’s visual engine, boasting strong artistic styling capabilities and precise language interpretation. Best for: Concept design and stylized artistic illustrations. |
3. Text Models
Text models can help develop prompts, structure ideas, analyze inputs, and support text-based steps inside a workflow.
To add and configure text-based workflow steps, see Text Generation Nodes.
| Provider | Available Models |
|---|---|
| OpenAI | GPT 5.5, GPT 5.4, GPT 5.4 mini, GPT 5.4 nano, GPT 5.2, GPT 5.1, GPT 5 mini, GPT 5 nano |
| Anthropic | Opus 4.8, Opus 4.7, Opus 4.6, Sonnet 4.6, Haiku 4.5 |
| DeepSeek | DeepSeek V4 Pro, DeepSeek V4 Flash |