A multimodal video model with sound for rapid creative validation
Seedance 2.0 Fast is a fast variant of ByteDance's Seedance 2.0 series, suitable for short-film concepts, character shot tests, and advertising asset iteration. doubao-seedance-2-0-fast-260128 supports text, image, and audio/video references, and can generate videos with sound, delivering short clips in 480p or 720p while balancing character references and creative control.
Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.
Specifications and API features
Output resolution
480p, 720p
Video duration
4–15 seconds; duration=-1 uses automatic duration
Aspect ratio
16:9, 4:3, 1:1, 3:4, 9:16, 21:9, adaptive
Creation methods
Text-to-video, image-to-video with first frame and first/last frames, multimodal image and audio/video references
Reference assets
Up to 9 images, 3 audio clips, and 3 videos
Sound and camera control
generate_audio for sound generation; camerafixed for a fixed camera; seed for random seed
Invocation and delivery
POST /seedance/videos; supports asynchronous tasks, callbacks, and video link delivery
The above are the available specifications and control methods for this model on this platform; API parameters do not represent all native model capabilities.
Core capabilities
Preserve characters, redesign shots
Use reference_image to provide a character or subject appearance, then describe new scenes, clothing, actions, and camera movement in text. It is suitable for creations where “the same person enters different stories,” without treating the photo background as a fixed scene; clear, front-facing, unobstructed character assets make it easier to establish references.
Let motion and sound participate in creation
In addition to images, you can add reference_video and reference_audio to guide actions, camera movement, and sound rhythm. Enable generate_audio to generate video with sound; prompts should separately explain what happens visually, how the camera moves, and what you want to hear, avoiding conflicting intentions across multiple assets.
Switch between composition control and flexible references
Use first_frame when motion needs to begin from a specified image, and add last_frame when you need to control the ending image; use reference mode when you want to redesign the composition. You can also reduce camera movement changes with a fixed camera and request the final frame to check how the shot concludes or prepare the starting image for the next segment.
Use Cases
Character Short Film Storyboard Prototyping
Upload authorized character photos and a single-shot description, such as a person walking in a park and turning to smile, then select a portrait aspect ratio to generate a short video. Keep the reference character unchanged while modifying the scene and actions separately to deliver a set of candidate storyboards for selecting character performance and narrative direction.
Product Advertising Shot Iteration
Use a product image as the first frame, specify a camera push-in, lighting changes, or environmental motion, validate the concept in 480p first, then generate the selected version in 720p. If the background needs to be rearranged, use the subject reference mode instead; the deliverable is short shots for editing selection, not a complete advertising project.
Concept Previsualization for Scenes with Audio
Write scenes such as rainy streets, beaches, or indoor activities as prompts, and use legal audio and video materials to guide the rhythm while enabling audio generation. Output audiovisual drafts in landscape or widescreen aspect ratios for directors, animation teams, or content editors to discuss environmental atmosphere, motion direction, and sound design.
How to Choose This Model
Choose Fast for Rapid Prototyping, the Standard Version for High Resolution
When the delivery target is 480p or 720p short shots and character and audio-video references are needed, 2.0 Fast is suitable for creative prototyping and multi-option selection. If the task explicitly requires 1080p or 4k, choose the 2.0 standard version; Mini is a lightweight option in the same series. Fast's positioning does not mean every task has a fixed completion time.
Choose Short-Shot Generation and Long-Form Video Editing Separately
2.0 Fast is suitable for generating new 4–15-second shots and reference videos should not be understood as a full video editing mode. If you need up to 30 seconds, creation based only on audio references, or explicit editing and extension of existing videos, choose Seedance 2.5. The selection should be based on the required operation, not just the model generation.
Get Started
Organize content and Assets
In content, assign specific roles to images, videos, and audio, and explain the purpose of each asset in the text; organize first-and-last-frame constraints and subject references separately according to the task.
Select the Exact Version and Shot Settings
Specify model=doubao-seedance-2-0-fast-260128 for /seedance/videos; first test with duration=5, resolution=720p, and a clear aspect ratio. When audio is needed, explicitly set generate_audio=true; it is disabled by default.
Save the Task and Final Frame
For asynchronous processing, first obtain the task_id, query /seedance/tasks or receive a callback; after completion, check the subject, actions, ending, and audio, then save the finished video and returned final frame as needed.
Trial recommendation: Ad action preview
Input and goal
The character reference image defines the clothing, the video reference defines a slow turn, and the character completes the turn in a bright studio while the camera remains steady.
Acceptance and next steps
Use a 720p short shot to verify the subject and action, then arrange versions with and without sound; do not use standard edition 4k settings for Fast.
Usage limits
First and last frames and multimodal references cannot be mixed in the same request. Use first_frame and last_frame when strictly specifying the start and end images; in reference mode, text can specify which image serves as the beginning or end, but this guidance is not equivalent to strictly locking the first and last frames.
Reference audio must be wav or mp3, 2–15 seconds per clip, no more than 15MB per clip, and no more than 15 seconds in total; reference video must be mp4 or mov, 2–15 seconds per clip, with no more than 15 seconds in total. Assets outside these limits may fail during processing.
This version does not provide 1080p or 4k output, and does not use the 2.5-exclusive edit, extend, or output format selection. Human and character assets must be owned or authorized; character references are used to guide appearance and should not be considered a guarantee of precisely reproducing identity details in every frame.
Frequently asked questions
What is the difference between a reference image and a first-frame image?
reference_image is used to guide the appearance of a character or subject, while the scene, clothing, and action can be described anew; first_frame makes the video begin moving from the given image. Choose the first frame if you want to preserve the original image composition, and choose a reference image if you want the same character to enter a new scene. Do not mix the two modes.
Can 2.0 Fast generate videos with sound?
Yes, enable it by setting generate_audio=true; audio is not generated by default. Prompts can describe ambient sounds or action sounds; reference_audio is used as a reference for sound and rhythm, and does not mean the final video will necessarily reproduce the input audio word for word or segment by segment.
Can it generate 30-second or 4k videos?
This version uses 480p, 720p, and a 4–15 second range, and automatic duration can also be selected. For 30-second tasks, choose Seedance 2.5; for 4k tasks, choose the 2.0 standard edition. Automatic duration lets the model determine the length; it does not remove this model's duration limit.
Can I upload a video for Fast to directly edit or extend?
Fast can use a video as reference_video to help guide the action and camera movement of a new shot, but do not treat it as an edit or extend operation. When you need to explicitly modify an existing video or extend a clip, choose 2.5 and organize assets and parameters according to the corresponding task mode.
How do I submit a task and retrieve the generated video?
Submit the exact model value and content array to /seedance/videos. Text items can contain up to 1000 characters, and image addresses use image_url objects. You can set async=true and then query task_id, or configure callback_url to receive completion notifications, then retrieve the video from data.video_url.