All models

omnihuman-1.5

DreaminaVideo
1.46 Credits
Get your API key
omnihuman-1.5

Generate natural lip-sync digital humans from a single portrait and voice

OmniHuman 1.5 is an audio-driven video model for character narration: starting with a single character photo and a voice recording, it makes static images speak while coordinating lip movements, expressions, and head-and-shoulder movements. It is suitable for product explanations, course introductions, and brand broadcasts. You can adjust emotion and style through prompts, and use a subject mask to specify the appearing person in multi-person photos.

DreaminaModel brand
VideoModel type
VideoTask capability

Specifications and API features

Clarify capacity, inputs and outputs, and invocation methods before selecting a model.

Creation method
Generate a narration video from a single character image + driving audio
Input assets
Publicly accessible image URLs and mp3/wav audio URLs
Audio recommendation
Recommended to keep within 60 seconds; this is not a hard duration limit
Performance control
Use prompts to adjust expressions, emotion, stability, and style
Subject selection
mask_url supports specifying the driven person in multi-person photos
Tasks and output
Synchronous or asynchronous generation, returning task status and video_url

The above are this platform's asset requirements and invocation methods; the recommended audio duration does not represent the model's native limit.

Core capabilities

Learn what omnihuman-1.5 can bring to your work.

Voice drives complete narration performance

OmniHuman 1.5 does more than animate the mouth in a photo; it also coordinates expressions and head and shoulder movements with the rhythm of speech. The creative focus is how the character presents the content, making it suitable for narration videos that require a consistent presenter. When preparing audio, first determine pauses, emphasis, and the overall delivery tone.

Use prompts to shape the delivery tone

In addition to photos and voice, you can use prompts to describe emotions such as gentleness, calmness, and naturalness, as well as request subtle head movements or stable performance. This guides the character's performance style and is suitable for creating course explanations, brand broadcasts, and warm content separately, rather than replacing the lines in the audio.

From character assets to downloadable video

Use a clear front-facing portrait as the visual foundation, without preparing multi-angle modeling assets in advance. Multi-person photos can be paired with a subject mask to specify the driven subject; after generation, you receive a video address to continue downloading, editing, or adding to product pages, allowing character narration to integrate into existing content production workflows.

Use Cases

Start with specific tasks to identify where the model can be effective.

Product Updates and Feature Explanations

Turn product updates into voiceovers, pair them with a fixed presenter photo, and generate spoken video clips introducing new features. Then combine them with screen recordings, subtitles, and key annotations to deliver videos for help centers or new-user onboarding, with the person handling narration and the screen recording demonstrating actual operations.

Course Introductions and Training Materials

Record course openings, knowledge-point overviews, or training reminders as audio, pair them with instructor photos, and generate chapter-based digital human presentations. Using the same image maintains visual continuity across the course; when formulas, charts, or steps are involved, add courseware visuals in post-production rather than having the person present all information.

Brand Presentations and Content Reuse

Prepare different audio tracks for event announcements, product introductions, or frequently asked questions, and use the same brand persona to generate separate videos. Deliverables can further include product footage, subtitles, and end cards; reusing photos can reduce repeated filming, but each piece of content should still be checked to ensure expressions, lip movements, and audio match.

How to Choose This Model

Choose based on task complexity, input materials, and expected results.

Choose It When Photos and Voiceovers Are Ready

If you already have usable photos of a person and finished voiceover audio, and the goal is to have the person speak directly to the audience, OmniHuman 1.5's single-image audio workflow is straightforward. It is better suited to performances organized around sound rather than building complex narratives from text. Choose 1.5 based on presentation needs and sample results; there is no need to interpret the version number as a comprehensive improvement for every task.

Distinguish Presentations from Complex Scene Animation

For character explanations, course hosting, and brand announcements, consider OmniHuman 1.5 first; if the focus is on large movements, character blocking, scene changes, or precisely recreating a movement, consider animation or video models that support the corresponding controls. For presentation tasks, first test photos, audio, and prompts with short clips before producing complete content.

Get Started

From a small-scale task to formal integration.

01

Prepare the Task and Materials

Define the goal, required inputs, and output requirements, and start with real business examples.

02

Try It in the API Playground

Open the trial page, confirm the parameters supported by this entry point, then submit a small-scale task to review the results.

03

Integrate According to the API Documentation

Keep the complete model ID, use the request format specified in the documentation, and confirm billing rules on the Pricing page.

Usage Boundaries

Before official use, understand the output quality and scope of capabilities.

  • Portrait materials affect lip-sync and expression quality. Prioritize photos with even lighting, clear frontal views, unobstructed faces, and appropriate face proportions; profile, blurry, or overly dark images are not suitable for direct use as final production materials. For group photos, clearly specify the target person; designating a subject cannot be interpreted as multiple people speaking simultaneously.
  • Driving audio must be prepared in advance; a text script cannot replace the required audio material. Both image and audio URLs must be publicly accessible, and audio should be in mp3/wav format, preferably within 60 seconds. Longer content can be split into chapters for production, making it easier to check lip-sync and performance effects segment by segment before post-production stitching.
  • Prompts are used to guide emotion and style; they are not tools for frame-by-frame action choreography, nor do they guarantee identical performances every time. After output, check lip-sync, character movements, and content expression, and save the video promptly; when using real people's likenesses and voices, obtain the appropriate authorization first.

Frequently Asked Questions

Answers to common questions about using omnihuman-1.5.

Can OmniHuman 1.5 generate a talking video from text only?

This workflow requires a character image and driving audio; text alone cannot replace these two assets. You can first record or create the script as an mp3/wav, then submit the audio URL. The prompt is used to adjust expressions, emotion, and style; it does not directly read the lines aloud.

What conditions should the photo meet?

It is recommended to use a clear, front-facing portrait with good lighting and no obstructions, with the person taking up an appropriate portion of the frame. The photo needs a publicly accessible URL. Before formal production, you can test the photo with a short audio clip to check lip sync, expressions, and head-and-shoulder movement, then continue using the asset.

Can I make a specific person in a group photo speak?

You can submit an array of subject mask URLs through mask_url to specify the person to be driven in a multi-person photo; even if there is only one mask URL, it should still be submitted as an array. You should first identify the target and prepare the corresponding mask. This feature is for subject selection and should not be understood as allowing multiple characters to speak separately or complete a multi-person conversation in a single request.

How can I control tone and character movements?

The voice itself provides speaking speed, pauses, and delivery rhythm, while the prompt can add requirements such as gentle, calm, natural, or slight head movements. It is recommended to keep the voice-over and prompt consistent, avoiding an impassioned voice with text requesting calmness; the final performance still needs to be checked through the generated video.

How do I submit and retrieve videos in the application?

Submit image_url and audio_url to POST /dreamina/videos, and specify omnihuman-1.5. Results can be obtained synchronously for short tasks; for longer tasks, you can use callback_url, or set async:true and then query /dreamina/tasks. After completion, read video_url and save it.

Model information · Updated: 2026-10-01. For calling parameters and billing rules, see the API and pricing sections.

Use omnihuman-1.5 for your next task

Start with a clear goal and assess whether it suits your work based on actual results.