Create 480P and 768P short videos with text and reference materials
MiniMax-H3-Max is an audio-video model in the H3 series for rapid generation and iteration, supporting creation with text, first and last frames, and multimodal references. It uses unified content input to generate 5–15 second clips in 480P or 768P, making it suitable for advertising drafts, character shots, and motion concept comparisons; when 2K delivery is needed, compare MiniMax-H3.
Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.
Specifications and API features
Model ID
MiniMax-H3-Max
API endpoint
POST /minimax/videos
Input
Text, images, video, or audio in content; see the API documentation for combination requirements
Output resolutions
480P, 768P
Duration
Integer values of 5–15 seconds
Delivery
Asynchronous task_id, task queries, and optional callbacks
Core capabilities
Organize shots and sound together
Text can describe visual action, dialogue, sound effects, and atmosphere at the same time, while image, video, and audio references help convey creative intent. After generation, check not only the visuals but also listen to the native sound and lip-sync.
Choose first/last-frame and reference modes separately
Use first_frame/last_frame to define the starting and ending visuals, or use reference_image/video/audio to express subject, action, and sound references. The two material modes are mutually exclusive, so do not mix all roles in the same request.
Built for rapid iteration across multiple concepts
Max is officially positioned for rapid generation, and 480P/768P are suitable for comparing composition and action first. Choose integer durations of 5–15 seconds, evaluate speed and quality with real tasks, and do not treat the model name as a fixed completion time.
Use cases
Comparing advertising action concepts
Use the same product and character materials to try walking, turning, or demonstration actions separately. First select clips that communicate the selling points, then add brand copy and editing.
Character and motion references
Character images define appearance, video defines motion, and audio defines rhythm; explain the reference relationships in text, and check whether clothing, hands, and rhythm remain consistent.
First-and-last-frame transition previews
Prepare start and end images with consistent viewpoints, describe the action that happens in between, and generate a connecting shot; image constraints are not equivalent to frame-by-frame animation control, so the middle still needs to be checked.
How to choose this model
First compare speed and available actions
When deliverables are mainly 480P/768P clips, creative previews, or multi-option iterations, start with Max. It is a different model from MiniMax-H3, still uses the same content task format, and requires the full ID when integrating.
Choose H3 for 2K and a four-second starting point
Max supports 5–15 seconds and 480P/768P; H3 supports 4–15 seconds and 768P/2K. First choose based on duration, delivery resolution, and reference assets, then compare finished videos using real examples.
Get started
Use content to describe shots and references
At least one non-empty type=text entry is required; set images, videos, and audio by role. First-and-last-frame scenarios and reference asset scenarios are mutually exclusive, so do not submit both sets of controls at the same time.
Select the output range for this model
Set model=MiniMax-H3-Max for /minimax/videos, and choose 480P/768P and an integer duration of 5–15 seconds. Explicitly select an aspect ratio for text and reference modes; for first-and-last-frame mode, use adaptive based on the input images.
Receive task content.url
After asynchronously saving task_id, query /minimax/tasks or receive a callback, wait for task.status=succeeded, then read task.content.url; check the audio, visuals, and duration, then download it for editing.
Show the same product reference under three different lighting directions, starting with one option: soft side lighting, the product rotating slowly, a simple background, and natural sound.
Acceptance and next steps
Explore options at 480P or 768P for 5–15 seconds, and check the subject and audio-visual quality; when 2K is needed, choose H3, which supports that tier, rather than submitting 2K to Max.
Usage limits
H3 Max does not support 2K, and its duration range is integer values from 5–15 seconds.
The API uses the content structure and does not accept legacy top-level prompt, image_urls, audio_urls, or messages fields.
First-and-last-frame and multimodal reference modes cannot be mixed. A total of up to 12 reference assets is allowed; prepare inputs according to the quantity, duration, size, and format limits for each type.
Accept audio and visuals together: dialogue, lip sync, actions, and pacing may require post-production revisions. Fees for reference assets and output videos are subject to the current Pricing rules.
Frequently Asked Questions
How do I select H3 Max?
Submit a request to POST /minimax/videos, explicitly set model=MiniMax-H3-Max, and provide non-empty text in content.
What resolutions and durations are supported?
480P, 768P, and integer durations of 5–15 seconds are supported; this model does not support 2K.
How do I provide reference materials?
Use the content items and role values defined in the documentation to indicate the first frame, last frame, or reference purpose, and comply with the quantity and combination limits for each type of material.
How are reference materials billed?
Audio input is not charged extra, and the first two images are free; additional images and reference videos are billed based on their corresponding input usage, while output video costs are calculated according to the current rules.
How do I retrieve the completed video?
An asynchronous request first returns task_id. Query the same task, and retrieve task.content.url after confirming task.status=succeeded; read the task error information if it fails.