All models

happyhorse-1.0-i2v

HappyHorseVideo
Get your API key
happyhorse-1.0-i2v

Create natural, fluid dynamic clips starting from a first-frame image

happyhorse-1.0-i2v is HappyHorse 1.0's first-frame image-to-video model, using an image to define the starting point of the scene and then combining it with a text description to generate dynamic clips. It focuses on realistic dynamic rendering and image-text semantic understanding, making it suitable for turning product images, character concept art, or scene visuals into video assets. When creating, first establish the composition, then describe the subject's actions, environmental changes, and camera movement.

HappyHorseModel brand
VideoModel type
First-frame image-to-videoCreation method
STANDARD APIs · QUICK SETUP

Bring this model into your workflow

Submit requests to the public API at api.acedata.cloud using the documented parameters, then use the results in your application.

API hostapi.acedata.cloud
modelhappyhorse-1.0-i2v

Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.

Specifications and API Features

Creation method
First-frame image-to-video; input an image and text, output a video
First-frame input
image_url; use a publicly accessible image URL
Platform video duration
3–15 seconds, duration defaults to 5 seconds
Platform output resolution
720P / 1080P, default 1080P
Aspect ratio handling
Follows the first-frame image aspect ratio whenever possible; no ratio is required for image-to-video
Invocation method
POST /happyhorse/videos; action=image_to_video
Task delivery
Supports asynchronous queries and callbacks; returns video_url upon completion

Its native capability is positioned as video generation from image and text input; the duration, resolution, and task delivery methods here correspond to this platform's supported invocation scope.

Core Capabilities

Create from an established composition

The input image serves as the first frame of the video, allowing creation to begin with an already defined subject, background, and visual tone rather than rebuilding the entire scene with text. It is suitable for animating existing design assets; prepare an appropriate first frame first, then use prompts to supplement subsequent actions, making camera intent clearer.

Describe dynamic changes with text

The model combines image content and text semantics to generate video, emphasizing realistic motion and smooth, natural visuals. Prompts can be structured around what the subject does, how the environment changes, and how the camera moves, such as a person slowly turning their head, clothing swaying in the wind, or the camera moving forward. Avoid stacking contradictory actions.

Integrate asynchronous asset production

Generation tasks can be submitted using async, and callback_url can also be configured to receive completed results. After an application saves task_id, it can retrieve the status and video link through task queries without requiring the user interface to keep waiting for the same request. Results also include information such as duration and resolution, facilitating asset archiving and subsequent processing.

Use Cases

Animating Static Product Images

Input a product image with a complete composition, describe a subtle camera push-in, dynamic background, or display action, and generate short video candidates for product presentation. First use 720P to validate the motion direction, then choose 1080P as needed for delivery; inspect label text, edges, and key visual details frame by frame before proceeding to editing and publishing.

Character Shot Previsualization

Use a character design image or completed storyboard image as the first frame, supplement it with descriptions of actions such as turning, looking up, or walking, and create single-shot previsualization material. It is suitable for validating the visual effect of a static design after it enters motion, then allowing creators to select the action rhythm and camera direction, rather than replacing complete multi-shot continuity production.

Atmospheric Scene Clips

Use indoor, architectural, or natural scene images as a starting point, describe clouds, wind, light and shadow, and slow camera movement to create atmospheric clips. When you want vertical or square videos, first prepare the first-frame image according to the target composition; the output aspect ratio will follow the image as closely as possible, so you should not rely on the ratio parameter to reorganize the frame.

How to Choose This Model

If You Already Have a First Frame, Choose I2V

If the product, person, or scene is already defined in an image, choose happyhorse-1.0-i2v to start dynamic generation from the first frame. If you only have a text concept, choose the corresponding t2v model; if multiple images are needed to jointly constrain the character or style, choose the r2v model; if an existing video needs modification, use video-edit rather than only changing the action of this model.

Clearly Distinguish Between 1.0 and 1.1

happyhorse-1.0-i2v and happyhorse-1.1-i2v are different version call IDs. If you need to continue using 1.0 project configurations or compare against existing assets, explicitly specify this model; when trying 1.1, it is recommended to compare motion, details, and usable clips using the same first frame and prompt before deciding which version to use, rather than judging results by the version number alone.

Get Started

Prepare Inputs for the Corresponding Operation

Prepare a single first-frame image_url and use prompt to describe the action; keep the aspect ratio as close as possible to the input image, with no need to pass ratio.

Explicitly Specify the Version

Set model=happyhorse-1.0-i2v and action=image_to_video for /happyhorse/videos. Choose an integer duration of 3–15 seconds and 720P or 1080P, using the input image to determine the starting composition.

Query and Save Completed Videos

Use async or callback_url to integrate background tasks, save task_id and query /happyhorse/tasks; wait for succeeded before reading video_url, then check the person, motion, and audio.

Trial recommendation: Keep the main image's short animation

Input and goal

Use the product image as the first frame, retain the bottle shape and packaging, let the background fabric sway gently in a breeze, and keep the camera stable without cuts.

Acceptance criteria and next steps

Use image_to_video and image_url, and make the aspect ratio follow the input image as closely as possible; do not use ratio to forcibly alter the original composition.

Usage boundaries

  • This model generates video from a first-frame image; it is not for multi-reference image generation or editing existing videos. image_url determines the video starting point; image_urls, video_url, and settings for preserving original video audio should not be regarded directly as creative capabilities of this model.
  • The first-frame image must be accessible through a public URL, so ensure the link is reachable when submitted. Image-to-video output follows the input image ratio as closely as possible, so the target aspect ratio should ideally be finalized when preparing the image; do not interpret the general ratio option as forcing a new aspect ratio or recomposition.
  • Each generation is suitable for short clips of 3–15 seconds, not full-length video production. For material with complex occlusion, large movements, or text details that must be preserved precisely, first generate a small-scale test and check the result; voice-over, subtitles, and transitions between shots can be completed in post-production.

Frequently asked questions

How can I ensure I am calling 1.0 image-to-video?

When submitting a request to /happyhorse/videos, explicitly set model to happyhorse-1.0-i2v, action to image_to_video, and provide image_url. This directly specifies the version and task type, avoiding reliance on default values that may enter other generation modes.

Is the image the first frame or a regular reference image?

In this model's image_to_video action, image_url corresponds to the video's first frame and determines the initial composition. It differs from multi-reference image generation; if multiple character or style images need to participate in creation together, choose the corresponding r2v model.

Do I need to write a prompt?

This action requires an image, while text can describe the desired motion. Describe the subject's action, environmental changes, and camera approach rather than repeatedly listing every element in the image. Start by testing one clear action, then gradually add details, making it easier to evaluate the generated result.

Can it generate vertical videos?

Prepare a first-frame image with a vertical composition, then submit an image-to-video task. The output aspect ratio will follow the image as closely as possible, so there is no need to pass ratio separately; if the position of people, products, or backgrounds is important, arrange the composition and negative space in the first frame beforehand.

How do I retrieve the video after generation?

After submitting with async, save the task_id, then query the task through /happyhorse/tasks; you can also use callback_url to receive completion notifications. Only after the task completes successfully should you use the returned video_url to retrieve the video; do not treat a successful submission as meaning the video has already been generated.