Text-to-video and first/last-frame video model for short-shot creation
kling-v2-5-turbo is Kuaishou Kling V2.5 Turbo video model, suitable for turning text concepts or static images into short shots. It offers two main creation methods: text-to-video and image-to-video, and in pro mode, first and last frames can be used to arrange the starting and ending visuals. For product showcases, character motion, and storyboard trials, generation tasks can be organized around clear action and visual goals.
Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.
std, pro; the first frame guides image-to-video, and pro can be used with a last frame
Prompt control
cfg_scale range 0–1; negative_prompt up to 200 characters
Delivery format
Video URL, video ID, task ID, and status; supports asynchronous processing and callbacks
Talking photo workflow
Image and existing audio input; supports 5 or 10 seconds, combining animation and lip-sync
The above durations and control ranges apply to this platform's access endpoint; talking photo is a combined creation workflow and does not mean the model natively generates accompanying audio.
Core Capabilities
Structure a Complete Short Shot with Text
Use text2video to describe the subject, scene, action, and lighting, turning shot concepts into video. When creating, focus on one main event, such as a person turning around or a product entering the frame, then use negative prompts to exclude unwanted elements, making it easier to compare visual approaches created by different descriptions.
Arrange Visual Changes with Start and End Frames
Image-to-video uses the first frame as the visual starting point, making it suitable for projects with existing product images, character images, or storyboard frames. Pro mode can also add an end frame, giving the creation both a starting point and an end goal. Start and end frames guide generation rather than enabling frame-by-frame editing, so intermediate actions still need to be clearly expressed through prompts.
Match Character Photos to Existing Dialogue
In the photo talking workflow, submit a character image and prepared audio to create a dynamic talking-head video. V2.5 Turbo animates the photo, then completes lip-syncing based on the audio; prompts can describe actions and expressions during the animation stage, and the result includes both the final video and a link to the intermediate animated video.
Use Cases
Product Clips and Showcase Assets
Upload a product image as the first frame, describe the desired action, environment, and lighting, and generate short shots for product pages or promotional edits. When an ending composition is needed, choose pro and add an end frame; first determine whether the output is intended for landscape, portrait, or square use, then organize the assets to reduce composition loss caused by later cropping.
Storyboard Tests and Creative Comparisons
Break an individual shot from a script into descriptions of the subject, action, and scene, then use text to generate different visual versions; when storyboard images already exist, use the first frame to begin. The deliverables are video clips that can be viewed and edited, suitable for comparing shot direction and visual atmosphere rather than combining an entire multi-scene narrative into a single generation.
Photo Talking Videos and Character Introductions
Prepare a clear front-facing photo of one person and recorded dialogue audio, then use /kling/talking-photo to generate a character introduction or short talking video. The audio length is recommended not to exceed the selected video duration, and action prompts should mainly describe expressions and subtle movements; once complete, you can obtain both the lip-synced final video and the non-lip-synced animation clip.
How to Choose This Model
Choose pro When You Need Start and End Frames
If the goal is to create a short shot starting from one image, choose std or pro according to the task; when the ending image must also guide the result, choose pro. Compared with kling-v2-1-master, V2.5 Turbo's pro provides a start-and-end-frame creation method, making it better suited for tasks that already have starting and ending storyboards prepared, but this does not mean the intermediate process can be precisely locked frame by frame.
Choose a Version Based on Audio, Duration, and Editing Needs
The focus of V2.5 Turbo is text- or image-driven short video creation. If synchronized audio generation is needed, consider kling-v2-6's pro; if you need a duration of 3 to 15 seconds or 4K mode, consider kling-v3. If the task involves multi-image references or directly editing an existing video, choose kling-o1 or kling-v3-omni instead of adding these media fields to this model.
Get Started
Plan Text-Based Shots or Image Animation
Use text2video for text generation; use image2video and start_image_url for image-to-video, and clearly state the action and image preservation requirements in the prompt.
Distinguish std Start Frames from pro Start and End Frames
Send model=kling-v2-5-turbo to /kling/videos, and choose 5 or 10 seconds; std uses a start frame, while pro can provide both a start frame and an end frame. The end frame cannot be submitted alone.
Check the Start and End Images Before Editing
Save the task_id asynchronously, retrieve the finished video through task queries or callbacks, and check the subject and transitions. Standard generation has no native audio, so voiceover and music should be added in post-production; save selected clips as editing assets.
Trial Suggestion: Start-to-End Transition for a Product Ad
Input and Goal
The start frame shows a desk lamp turned off, and the end frame shows the desk lamp turned on and illuminating a book, while keeping the items on the desk and the camera position unchanged.
Acceptance and Next Steps
Evaluate the lighting change using pro's 5-second start-and-end frames; sound will not be generated at the same time, and camera movement descriptions are not a camera_control parameter.
Usage Limitations
Standard video generation for this model does not support generate_audio or native 4K mode. When dialogue is needed, you can use existing audio to create a talking photo video, or add voice-over in post-production; talking photos and standard video generation are different workflows and cannot be substituted by enabling the audio parameter.
This model does not support camera_control for dedicated camera movement control. You may describe camera intent in the prompt, but do not treat it as a precisely executable camera path. Start and end frames are also only used together in pro image-to-video, and an end frame cannot be submitted alone.
Image-to-video requires a first-frame image link; talking photos require accessible image and audio links. Driving audio supports mp3, wav, m4a, and aac, with a maximum size of 5MB; portrait photos are recommended to be clear, front-facing images of a single person, and audio length should match the intended 5-second or 10-second output.
Frequently Asked Questions
How can I ensure V2.5 Turbo is used when making a request?
Explicitly specify model=kling-v2-5-turbo in the /kling/videos or /kling/talking-photo request; do not rely on the default model. For video creation, choose either text2video or image2video; image-to-video requires a first frame, while talking photos require an image and audio.
What is the difference between std and pro for start and end frames?
V2.5 Turbo std does not support end frames, while pro supports start-and-end-frame guidance. When you need to specify the ending shot, choose image2video and pro, and submit both start_image_url and end_image_url; submitting only an end frame is not valid, nor can an end frame replace a first frame.
Can V2.5 Turbo directly generate videos with sound?
Standard video generation does not support generate_audio. If you already have voice-over audio, you can use the talking photo feature to combine a portrait image and audio into a lip-sync video; this is not the model creating sound on its own. For video tasks that need synchronized audio generation, choose a Kling version that supports this capability.
Can I use multiple reference images or reference videos to edit the scene?
This model is suited for text, first-frame, and pro start-and-end-frame creation, and should not be submitted with multiple images or reference videos according to the Omni workflow. When you need to reference multiple subjects, borrow video characteristics, or modify an existing video, choose kling-o1 or kling-v3-omni and organize the task according to the corresponding asset rules.
How do I receive the completed generated video?
Completed results provide video_url, video_id, task_id, and status information. For batch or background tasks, you can use async=true to obtain a task ID and then query the result, or configure callback_url to receive completion notifications; talking photo results additionally provide source_video_url for the intermediate animation.