How should I choose between V2.1 Master and V2 Master?
V2.1 Master is better suited as an option focused on visual quality and consistency. It is recommended to compare the actual results of both using the same prompt or first frame, rather than looking only at the version names. Both are used here for 5- or 10-second short clips, which does not imply additional control capabilities.
Can image-to-video specify both a first frame and a last frame at the same time?
This model uses first-frame image-to-video: submit start_image_url and describe the subsequent actions with a prompt; end_image_url last-frame constraints are not supported. If you must specify the ending image, you can choose a model that supports first and last frames, such as V2.5 Turbo pro, and prepare assets according to the corresponding creation workflow.
What is the difference between regular videos and talking photos?
Regular videos generate dynamic visuals from text or a first-frame image; this model does not natively generate synchronized audio. Talking photos, however, require a photo and existing audio, generating a spoken video through animation and lip synchronization; the audio is an input asset, not a voice automatically created by the model.
How do I submit a video task for this model?
Call POST /kling/videos, explicitly set model to kling-v2-1-master, choose text2video or image2video, and submit a prompt with a duration of 5 or 10 seconds. Image-to-video also requires a first-frame image link; after completion, retrieve the generated video through video_url.
How can I track generation results when producing in batches?
You can set async=true, first save the returned task_id, then query task progress; you can also configure callback_url to receive completion notifications. After obtaining results, associate video links and statuses by task ID. Photo narration can also use the asynchronous approach and returns the final video along with intermediate animation links.