All models

happyhorse-1.0-r2v

HappyHorseVideo
Get your API key
happyhorse-1.0-r2v

A short video model that reliably renders subjects and scenes with multi-image references

HappyHorse-1.0-R2V is a video generation model for reference-image creation, combining the subjects and scenes in images with action intent in text to generate new clips. It is suitable for creative tasks with existing character images, product images, or visual designs, focusing on preserving reference relationships rather than simply animating a single image. On this platform, you can explicitly select this version to complete multi-image input, aspect ratio settings, and video result retrieval.

HappyHorseModel brand
VideoModel type
Reference-based generationCreation method
STANDARD APIs · QUICK SETUP

Bring this model into your workflow

Submit requests to the public API at api.acedata.cloud using the documented parameters, then use the results in your application.

API hostapi.acedata.cloud
modelhappyhorse-1.0-r2v

Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.

Specifications and API Features

Creation mode
Reference image to video; image and text input, video output
Number of reference images
1–9 images, submitted via image_urls
Output resolution
720P, 1080P
Platform generation duration
3–15 seconds
Platform aspect ratios
16:9, 9:16, 1:1, 4:3, 3:4
Reference labels
Images can be referenced in order using names such as character1 and character2
Task delivery
Supports asynchronous queries, completion callbacks, and video URL retrieval

Multi-image references and image-to-text generation are model capabilities; duration, aspect ratio, and task delivery methods are listed according to this platform's supported usage scope.

Core Capabilities

Ground subjects and scenes in references

Creation does not need to rely solely on text descriptions of appearance. Submit subject images, scene images, and action prompts together, and the model can generate clips around clear visual references, with its key strength being the consistency of subject and scene references. It is suitable for short videos that need to retain character recognizability, product form, or visual direction, but does not mean copying images pixel by pixel.

Use multiple images to organize creative intent

A single task can include 1–9 reference images, with images referenced in prompts in order using names such as character1 and character2. You can specify which image provides the subject, what scene or visual elements to use, and then describe the action and camera work, reducing ambiguity in creative relationships when using multiple image inputs.

Plan generation around short-video delivery

The platform offers generation durations of 3–15 seconds, 720P and 1080P resolution, and aspect ratio options including landscape, portrait, and square. Tasks can be submitted asynchronously, with results retrieved through queries or callbacks, making it suitable for integrating reference-image creation into asset management, content review, and download workflows without having to wait continuously for requests to finish.

Use Cases

Product Showcase Clips

Provide product reference images and display environment images, and use text to describe product placement, camera movement, and visual atmosphere to generate short clips suitable for marketing proposal previews. Reference images convey appearance, while prompts organize actions; before delivery, focus on checking logos, structure, and small text, and avoid treating generated visuals directly as precise product records.

Character Storyboard Previsualization

Combine character design images with scene reference images, describe single-sequence actions such as a character entering the frame, turning, or walking, and generate storyboard previsualization materials. Suitable for discussing the relationship between characters and environments before filming or animation production; multiple shots can be generated separately, then check whether character appearances connect, rather than assuming continuity is automatically maintained across tasks.

Multi-Format Assets for the Same Theme

Using the same set of reference images, create horizontal, vertical, or square compositions separately to generate short video candidates for different display placements. Each prompt should specify the subject position and camera focus to make aspect-ratio trade-offs easier to compare. Output videos can enter the editing workflow for further addition of subtitles, sound, and brand information.

How to Choose This Model

Multiple Image References or a Fixed First Frame

When you need to combine character, prop, and scene references, choose 1.0-r2v; if the key requirement is to make a specific image the first frame of the video, choose an i2v variant. For text-only concepts, consider t2v; if you already have a video and want to change outfits or modify the visual style, choose video-edit. They correspond to different input methods, so simply changing the action name does not make them the same capability.

Explicitly Choose 1.0 Rather Than Relying on the Default Version

1.0-r2v and 1.1-r2v are different model variants, and reference-image generation defaults to 1.1. If you already have prompts and material test sets built around 1.0, you can explicitly specify 1.0 to continue evaluation; when trying 1.1, it is recommended to use the same reference images and tasks to compare subject retention and camera expression, rather than directly interpreting version numbers as a fixed degree of quality improvement.

Get Started

Prepare Inputs for the Corresponding Operation

Prepare 1–9 ordered image_urls, and use character1, character2, and so on in the prompt to reference characters or subjects by sequence number.

Explicitly Specify the Version

Set model=happyhorse-1.0-r2v and action=reference_to_video for /happyhorse/videos. Choose an integer duration of 3–15 seconds and 720P or 1080P; text and reference-image generation can set ratio.

Query and Save Completed Videos

Use async or callback_url to connect to background tasks, save task_id and query /happyhorse/tasks; wait for succeeded before reading video_url, then check characters, actions, and audio.

Trial recommendation: multi-image character and clothing reference

Input and goal

Image one defines the character, image two defines the clothing. Generate a shot of character1 walking through a courtyard wearing the reference outfit, with natural lighting and style.

Acceptance criteria and next steps

Arrange 1–9 image_urls in order and reference them in the copy; check the character identity, clothing, and new scene, and do not treat all images as fixed first frames.

Usage limitations

  • Reference images are used to guide generation and do not mean that the subject, clothing, logos, or scene details will remain completely unchanged. If multiple images conflict with one another in appearance or style, first select the materials and specify the purpose of each image in the prompt; content with high requirements for product details and character recognizability needs the results checked segment by segment.
  • This model is for generating video from reference images, not an existing video editing entry point or a fixed first-frame animation entry point. Do not treat video_url or original audio preservation settings as its primary creative method; when you need to modify existing clips or preserve the original audio, choose the corresponding video editing model.
  • Platform-generated clips are 3–15 seconds long; long narratives need to be split into multiple tasks and edited afterward. Image URLs must be publicly accessible; local paths or links that require login are not suitable as material inputs. After submission, you must also distinguish between pending, successful, and error statuses, and cannot determine that the video has been generated based only on the task ID.

Frequently Asked Questions

What must be provided when calling 1.0-r2v?

Submit a request to /happyhorse/videos, explicitly set model to happyhorse-1.0-r2v and action to reference_to_video, and provide a prompt and 1–9 image_urls. Do not rely on default actions or default models, otherwise they may not match the reference-image generation task.

Can it be used with only one reference image?

Yes, the number of reference images can start from 1. One image is suitable for providing a subject or visual direction, while multiple images are suitable for supplementing different reference relationships. If you require this image to directly become the first frame, choose the i2v variant; R2V focuses on using reference information to create new clips.

How can prompts accurately reference multiple images?

In the order arranged in image_urls, use names such as character1 and character2 to refer to the corresponding images, and clearly specify each purpose. For example, state that the subject references the first image and visual elements reference the second image, then describe the action and camera. After adjusting the image order, you should also update the prompt accordingly.

Can dialogue be generated directly and the original video audio preserved?

Do not design tasks assuming that dialogue, lip-sync, or original audio preservation are confirmed features of 1.0-r2v. Its stated purpose is to generate video from image and text references; preserving existing video audio is an operation of video-edit. When specific audio content is required, voice-over and mixing can be completed in post-production.

How do I get the video after asynchronous submission?

After submitting with async, retain the task_id and query it through /happyhorse/tasks; you can also provide callback_url to receive the completion result. Task statuses include pending, succeeded, and error. After success, obtain video_url and use the returned duration and resolution to check whether the final video meets the requirements of this task.