No video yet. Submit the form to generate video.
Explore different use cases and parameter configurations
Transparent pricing with no hidden fees. Pay as you go.
Total Cost = Unit Price × (Generated Video Duration + Input Video Duration) + Additional Image Cost
| Rule & Modality | Channel | Credits | Price (USD) | Official / Reference Price | Daily Savings |
|---|---|---|---|---|---|
MiniMax-H3-video-input-768p videoMiniMaxresolution: 768P | Starter | 1per second | $0.005 | MiniMax$0.090per second | Based on 1,000 10-second videos/dayMiniMax costs $854.550/day moreApiPass is 94.95% lower |
MiniMax-H3-text-to-video-2K videoMiniMaxresolution: 2K | Starter | 1per second | $0.005 | MiniMax$0.140per second | Based on 1,000 10-second videos/dayMiniMax costs $1354.550/day moreApiPass is 96.75% lower |
Competitor pricing costs $1354.550/day more
ApiPass is 96.75% lower, estimated at Based on 1,000 10-second videos/day
MiniMax-H3-text-to-video-2K
resolution: 2K
ApiPass Price
$0.005
1 credits per second
Competitor Prices
Complete guide to using MiniMax H3
Use MiniMax H3 API on ApiPass to turn prompts, frame images, and multimodal references into polished videos with up to 2K resolution, native stereo sound, consistent motion, and precise creative control.

MiniMax H3, also known as Hailuo 03, is a multimodal video generation system that understands text, images, video, and audio within a unified creative context. It supports text-to-video, image-to-video with first and last frame control, and reference-to-video using image, video, and audio assets. Through ApiPass, developers submit the public model name minimax/hailuo-h3, and ApiPass automatically selects the appropriate generation mode from the media fields in the input object.
Generate complete audiovisual clips directly from natural language prompts. Describe environments, characters, camera movement, shot progression, visual style, dialogue, music, and ambience in one request.
Animate a single first frame or create a controlled transition between first and last frames. This workflow is useful for product showcases, key-art animation, animated posters, UI demonstrations, and structured visual transitions.
Combine prompts with reference images, video clips, and audio files. Assign each asset to roles such as identity, product appearance, motion, camera language, editing rhythm, voice, or sound design.
MiniMax H3 can interpret text, images, video clips, and audio recordings within the same creative context, helping characters, motion, style, and sound remain connected throughout the generated video.
MiniMax H3 generates up to 2K video with native stereo sound, allowing visual action, dialogue, music, ambience, and sound effects to be directed together instead of added in a separate step.
Reference-to-video can combine multiple images, videos, and audio files in one request. Natural-language instructions clarify how each asset should influence identity, movement, style, voice, or atmosphere.
MiniMax H3 follows detailed directions for shot order, timing, transitions, camera movement, preservation requirements, readable text, and localized changes while maintaining unaffected scene details.
Create product films with close-up feature shots, branded typography, synchronized sound, and consistent materials across changing camera angles and scene transitions.
Generate opening sequences, animated posters, music visuals, editorial motion graphics, geometric layouts, typography-driven sequences, and sound-aware visual transitions.
Build short dramas, character promos, motion comics, and stylized narratives with consistent faces, clothing, props, voices, and performance details.
Prototype interface animations, game menus, equipment customization sequences, cinematic gameplay concepts, and digital experiences that combine interface logic with animated storytelling.
Choose the best MiniMax H3 workflow based on the source materials available for your application.
Use text-to-video when you want to start from a written prompt only. Provide prompt, duration, resolution, and a fixed aspect ratio such as 16:9 or 9:16.
Use image-to-video when you have a first frame or both first and last frames. ApiPass determines output ratio from the first frame, so aspect_ratio is not forwarded in this mode.
Use reference-to-video when you want images, videos, or audio to guide identity, motion, camera language, style, or sound. Reference mode supports adaptive and fixed aspect ratios.
Current regular pricing is 6 credits per second for 768P and 10 credits per second for 2K. A 5-second 2K generation costs 50 credits before any reference-media additions.
Reference video duration is added to billable seconds. Reference images after the first 5 add 6.4 credits each, while reference audio URLs currently do not add extra ApiPass H3 charges.
Failed tasks are refunded according to the existing task settlement logic. Use channel auto for automatic routing, or choose supported channels such as regular when you need more direct control.
Explore text-to-video, image-to-video, and reference-to-video in the ApiPass Playground. Test prompts, durations, resolutions, aspect ratios, first and last frames, and multimodal references before writing code.
Send a POST request to /api/v1/jobs/createTask using the public model name minimax/hailuo-h3. Provide prompt inside input, choose duration and resolution, and add frame or reference media when needed.
Use callBackUrl for success or failure notifications, and keep /api/v1/jobs/recordInfo as a polling fallback. Read generated video URLs from resultJson.resultUrls when the task state becomes success.
Use publicly accessible media URLs, validate mode conflicts before submission, keep prompts within 2000 characters for stable routing, and set reasonable polling intervals for active task retrieval.
All APIs require authentication via Bearer Token.
Authorization: Bearer
Submit a MiniMax Hailuo H3 video generation task using text, frame images, or reference media.
The API accepts a JSON payload with the following structure:
1{
2 "model": "minimax/hailuo-h3",
3 "callBackUrl": "https://example.com/api/callback",
4 "channel": "auto",
5 "input": {
6 "prompt": "A cinematic ocean sunrise with realistic waves.",
7 "duration": 5,
8 "resolution": "2K",
9 "aspect_ratio": "16:9"
10 }
11}modelRequiredstringPublic ApiPass model name. Must be minimax/hailuo-h3; do not pass minimax-h3, hailuo-03, or ApiPass model names with generation modes.
"minimax/hailuo-h3"
inputRequiredobjectH3 model input parameters. ApiPass uses these fields to normalize the request and automatically select text-to-video, image-to-video, or reference-to-video mode.
{"prompt":"A cinematic ocean sunrise.","duration":5,"resolution":"2K","aspect_ratio":"16:9"}callBackUrlOptionalstringCallback URL for task success or failure notifications. The field name is case-sensitive and must be callBackUrl. If omitted, no callback will be sent.
"https://example.com/api/callback"
channelOptionalstringProvider and billing channel used by APIPASS for the task. default channel is auto; APIPASS automatically allocates tasks across available providers based on real-time pricing and stability metrics; starter is ultra-low-cost with limited quotas and weaker stability; regular is standard and much cheaper than official APIs with moderate stability; official uses the model native API with high stability, fast task execution, and official pricing.
The default value for the channel parameter is auto. When enabled, APIPASS automatically allocates tasks across available providers based on real-time pricing and stability metrics to balance minimal cost and reliable performance. starter is an ultra-low-cost tier with limited quotas and weaker stability; regular is the standard tier and much cheaper than official APIs with moderate stability; official uses the model native API with high stability, fast task execution, and official pricing.
Available options:
auto
fallback_modelsOptionalarrayOptional ApiPass model-level fallback list. Each fallback item must provide its own model, channel, and complete input object.
[{"model":"minimax/hailuo-h3","channel":"auto","input":{"prompt":"A cinematic ocean sunrise.","duration":5,"resolution":"2K","aspect_ratio":"16:9"}}]The input object contains the following parameters:
input.promptRequiredstringVideo generation prompt. It must be a non-empty string. ApiPass keeps up to 7000 characters; MInimax H3 API receives up to 7000 characters, while ApiPass receives only the first 2000 characters if the prompt is longer.
For cross-provider consistency, keep prompts within 2000 characters. In reference mode, refer to media by array order, such as Image 1, Video 1, or Audio 1.
"A cinematic ocean sunrise with realistic waves and smooth camera movement."
input.durationOptionalintegerOutput video duration in seconds. Public valid range is 4 to 15 inclusive, with a default of 5.
Invalid, missing, fractional, boolean, or out-of-range values are normalized to 5. ApiPass uses 5 seconds when the public request duration is 4.
Available options:
5
input.resolutionOptionalstringOutput video resolution. Valid public values are 768P and 2K. Missing or invalid values default to 2K. The compatible alias quality maps to resolution, but resolution takes priority when both are present.
Values are trimmed and normalized to uppercase, so 768p and 2k are accepted. ApiPass currently outputs 2K even when the public request uses 768P.
Available options:
"2K"
input.aspect_ratioOptionalstringOutput video aspect ratio. Fixed values are 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16. Reference-to-video also supports adaptive. Compatible aliases are aspectRatio and ratio, with priority aspect_ratio > aspectRatio > ratio.
Text-to-video defaults to 16:9 when missing or invalid. Image-to-video does not forward aspect_ratio because output ratio is determined by the first frame. Reference-to-video supports adaptive and fixed ratios.
Available options:
"16:9"
input.first_frame_urlOptionalstringFirst-frame media URL for image-to-video mode. Compatible aliases by priority are first_frame_url, firstFrameUrl, image_url, imageUrl, and image.
Accepted as a string, an empty array, or a single-element array. Empty arrays and blank strings are treated as not provided; arrays with more than one element cause a parameter error.
"https://example.com/first.png"
input.last_frame_urlOptionalstringOptional last-frame media URL for image-to-video mode. Compatible alias: lastFrameUrl.
Use last_frame_url together with first_frame_url for consistent behavior across providers. It follows the same string and single-element array compatibility rules as first_frame_url.
"https://example.com/last.png"
input.reference_image_urlsOptionalarray(string)Reference image URL array for reference-to-video mode. Maximum 9 elements. Empty strings are ignored during normalization. Compatible alias: referenceImageUrls.
Cannot be combined with first_frame_url, last_frame_url, or image_urls. Reference image URLs must be publicly accessible and should point to real media content.
["https://example.com/subject.png","https://example.com/environment.png"]
input.reference_video_urlsOptionalarray(string)Reference video URL array for reference-to-video mode. Maximum 3 elements. ApiPass requires each reference video to be 2–15 seconds and total reference video duration to be no more than 15 seconds. Compatible alias: referenceVideoUrls.
The URL must be publicly accessible over http or https and no longer than 2048 characters. ApiPass reads actual reference video duration during task creation for accurate billing and rejects the task if duration cannot be read.
["https://example.com/motion.mp4"]
input.reference_audio_urlsOptionalarray(string)Reference audio URL array for reference-to-video mode. Maximum 3 elements. Empty strings are ignored during normalization. ApiPass supports audio-only reference mode; ApiPass requires reference audio to be combined with reference images or videos. Compatible alias: referenceAudioUrls.
ApiPass does not support audio-only reference mode. If only reference audio is provided and the task routes to ApiPass, the adapter ignores the audio and falls back to text-to-video with a 16:9 ratio.
["https://example.com/atmosphere.mp3"]
1curl --location 'https://api.apipass.dev/api/v1/jobs/createTask' \
2 --header 'Authorization: Bearer YOUR_APIPASS_API_KEY' \
3 --header 'Content-Type: application/json' \
4 --data-raw '{
5 "model": "minimax/hailuo-h3",
6 "channel": "auto",
7 "input": {
8 "prompt": "A cinematic ocean sunrise with realistic waves and smooth camera movement.",
9 "duration": 5,
10 "resolution": "2K",
11 "aspect_ratio": "16:9"
12 }
13 }'1{
2 "code": 200,
3 "message": "success",
4 "data": {
5 "taskId": "task_0123456789abcdef"
6 }
7}codeStatus code. 200 indicates success; other values indicate failure.
messageResponse message. When the request fails, this contains the error description.
data.taskIdApiPass task ID used to query task status and results.

ByteDance
ByteDance
Starting from
0 credits
ByteDance
ByteDance
Starting from
0 credits
ByteDance
ByteDance
Turn text, images, video clips, and audio into cinematic, audio-synced videos with a single API call.
Starting from
0 credits
ByteDance
ByteDance
Generate cinematic 4K AI videos up to 30 seconds with Seedance 2.5's multimodal power.
Starting from
0 credits
Google Veo 3.1 is the latest state-of-the-art generative video model designed to transform text and image prompts into high-fidelity cinematic visuals. Building upon its predecessors, Veo 3.1 features significant upgrades in prompt adherence, visual realism, and native audio generation, creating 8-second clips with synchronized sound effects and dialogue.
Starting from
0 credits
wan
wan
Starting from
0 credits
Runway
Runway
Starting from
0 credits
Kling
Kling
Transform text prompts and images into high-quality videos with advanced AI. Generate cinematic content with customizable duration, aspect ratio, and audio capabilities.
Starting from
0 credits
MiniMax
MiniMax
Generate high-quality AI videos from text prompts or images with MiniMax's Hailuo 2.3 model. Support for both 6s and 10s durations, 768p and 1080p resolutions, with intelligent prompt optimization.
Starting from
0 credits
grok
grok
Grok Imagine is xAI’s AI video generation model for creating high-quality videos with realistic motion, native audio, and synchronized speech. Grok Imagine 1.5 introduces faster generation speeds, improved motion physics, and stronger visual consistency for modern AI video applications.
Starting from
0 credits
Luma
Luma
Starting from
0 credits
Kling
Kling
Kling v3 video generation model, supports text-to-video and image-to-video, up to 15 seconds, 1080p pro mode
Starting from
0 credits
wan
wan
Generate videos with cinematic consistency by reimagining existing footage through high-fidelity visual transformation and structural control.
Starting from
0 credits
Starting from
0 credits
grok
grok
Starting from
0 credits
kling
kling
Starting from
0 credits