No video yet. Submit the form to generate video.
Explore different use cases and parameter configurations
Transparent pricing with no hidden fees. Pay as you go.
| Rule & Modality | Channel | Credits | Price (USD) | Official / Reference Price | Daily Savings |
|---|---|---|---|---|---|
with-video-720p videoGoogleresolution: 720p | Regular | 268per run | $1.218 | - | - |
with-video-1080p videoGoogleresolution: 1080p | Regular | 270per run | $1.227 | - | - |
with-video-4k videoGoogleresolution: 4k | Regular | 400per run | $1.818 | - | - |
no-video-720p-4s videoGoogleduration: 4, resolution: 720p | Regular | 102per run | $0.464 | - | - |
no-video-720p-6s videoGoogleduration: 6, resolution: 720p | Regular | 135per run | $0.614 | - | - |
no-video-720p-8s videoGoogleduration: 8, resolution: 720p | Regular | 168per run | $0.764 | - | - |
no-video-720p-10s videoGoogleduration: 10, resolution: 720p | Regular | 200per run | $0.909 | - | - |
no-video-1080p-4s videoGoogleduration: 4, resolution: 1080p | Regular | 103per run | $0.468 | - | - |
no-video-1080p-6s videoGoogleduration: 6, resolution: 1080p | Regular | 136per run | $0.618 | - | - |
no-video-1080p-8s videoGoogleduration: 4, resolution: 1080p | Regular | 168per run | $0.764 | - | - |
no-video-1080p-10s videoGoogleduration: 10, resolution: 1080p | Regular | 202per run | $0.918 | - | - |
no-video-4k-4s videoGoogleduration: 4, resolution: 4k | Regular | 233per run | $1.059 | - | - |
no-video-4k-6s videoGoogleduration: 6, resolution: 4k | Regular | 268per run | $1.218 | - | - |
no-video-4k-8s videoGoogleduration: 8, resolution: 4k | Regular | 300per run | $1.364 | - | - |
no-video-4k-10s videoGoogleduration: 10, resolution: 4k | Regular | 335per run | $1.523 | - | - |
Complete guide to using Gemini Omni
Generate cinematic, physics-aware AI videos with Google's Gemini Omni model — combined with reusable characters and voices — through a single ApiPass endpoint.

The Gemini Omni Video API on ApiPass is the video-generation surface of Google's Gemini Omni Flash, a high-performance multimodal model designed for high-speed video generation, editing, and cinematic control. Instead of running text-to-video as an isolated call, ApiPass exposes Gemini Omni Video as the composition layer of the Omni family: you can pass in reusable character resources and voice profiles that you've already created on ApiPass, and get back a fully rendered video with synchronized visuals and audio. Gemini Omni processes text, image, audio, and video simultaneously, giving you more cohesive, consistent, and controllable output — and ApiPass wraps that capability behind one API key alongside the rest of its model catalog.
Instead of managing separate accounts, preview access, or cloud setup, ApiPass gives you Gemini Omni Video behind the same API key you already use for the rest of the model catalog.
Gemini Omni Video accepts any mix of text, image, video, and audio in a single prompt, letting you seed generations from real creative materials rather than a blank text prompt.
Grounded in a model of real-world physics, the API renders believable reflections, gravity, lighting, and weather, so scenes hold together visually instead of drifting into artifacts, even in dynamic shots.
You can guide a generation with multiple reference images and short video clips, and the model keeps subjects, style, and scene consistent, holding identity steady across edits and shots.
Conversational editing lets you iteratively refine and edit your videos through natural language, so you can rework a scene without rebuilding the entire prompt from scratch.
Gemini Omni combines broad world knowledge with strong video capabilities — it understands cultural context, historical settings, and scientific accuracy, generating videos that make sense, not just look pretty, which is useful when prompts reference specific eras, locations, or concepts that require factual coherence.
Gemini Omni API supports natural language video editing, allowing users to refine a scene step by step instead of rebuilding the full prompt each time. A user can change the environment, adjust the action, replace objects, shift the camera angle, or add visual effects while keeping the original scene coherent. This makes it useful for AI video editors, creator tools, and applications where users need a more intuitive way to transform existing footage.
Gemini Omni Video connects visual creation with knowledge of physics, history, biology, culture, and narrative logic. This helps generated videos feel less random and more intentional, which is valuable for explainers, cinematic storytelling, product concepts, and educational content.
Text can define direction, images can guide subjects or style, video can provide motion and scene context, and audio or character resources can support speech-driven or identity-aware workflows. This helps developers build video tools that start from real creative materials instead of a blank prompt alone.
Gemini Omni Video can support avatar-style scenarios where character presence, expression, delivery, and environment need to feel connected. This is useful for presenter clips, character-led content, interactive media, and future-facing creative video products.
Vertical clips generated from a prompt plus a single reference image, with a repeatable protagonist across every episode.
Localized ad variants that share one on-brand spokesperson and voice but swap product, message, or market.
Educational content, historically accurate scenes, and knowledge-rich storytelling where factual coherence matters as much as visual quality.
Speech-driven, avatar-style workflows where vocal delivery guides the final result, useful for presenter videos, character dialogue, narrated scenes, and generated clips where voice, expression, and on-screen action need to feel connected.
Official Gemini Omni access is currently split across preview tracks and multiple consoles and enterprise surfaces. ApiPass consolidates it behind a single API key alongside the rest of your model stack, so integration is a base-URL and key change rather than a new provider onboarding.
The official API treats each generation as a self-contained request. On ApiPass, Gemini Omni Video is designed to consume the character and voice resources you've already built on the platform via the Gemini Omni Character and Gemini Omni Audio capabilities — so you can maintain a library of identities and voices and plug them into any subsequent video job.
Producing Episodic Short-Form Video That Needs A Repeatable Protagonist, Narrator, And Visual Style.
Shipping Localized, Personalized Ad Variants That Stay On-Brand Across Markets And Creative Permutations.
Turning Course Scripts And Lesson Plans Into Narrated, Character-Led Instructional Videos With Historically Or Scientifically Grounded Scenes.
Producing Cinematics, Trailers, And In-App Story Sequences Where The Same Character Reliably Appears Across Many Shots And Iterations.
Create an ApiPass account and generate an API key from your dashboard to unlock the full Gemini Omni family — video, character, and audio — alongside the rest of the model catalog.
Write a scene prompt and optionally gather your reference materials — images, a short source video, a reusable character resource, and a voice profile from Gemini Omni Audio — so the model has everything it needs to keep identity, tone, and style consistent.
Call Gemini Omni Video with your prompt and chosen inputs, review the generated clip, and refine it through conversational edits until the scene, motion, and delivery match your creative direction.
All APIs require authentication via Bearer Token.
Authorization: Bearer
Submit a new Nano Banana 2 image generation or editing task
The API accepts a JSON payload with the following structure:
1{
2 "model": "string",
3 "callBackUrl": "string (optional)",
4 "channel": "auto",
5 "input": {
6 // Input parameters
7 }
8}modelRequiredstringThe model name to use for generation
"google/nano-banana-2"
callBackUrlOptionalstringCallback URL for task completion notifications. If omitted, no callback will be sent.
"https://your-domain.com/api/callback"
channelOptionalstringYou may specify the corresponding provider within APIPASS via the channel parameter; these providers handle the actual image and video generation tasks. APIPASS currently offers three provider options:
The default value for the channel parameter is auto. When enabled, APIPASS automatically allocates tasks across available providers based on real-time pricing and stability metrics to balance minimal cost and reliable performance. Retain the default auto value unless you have custom routing requirements.
Available options:
auto
The input object contains the following parameters:
input.promptRequiredstringA text description of the image you want to generate
Describe the subject, style, lighting, and composition for best results
"A serene alpine lake reflecting snow-capped mountains at golden hour, photorealistic"
input.image_inputOptionalarray(URL)Input images to transform or use as reference. Supports up to 14 images.
Accepted types: image/jpeg, image/png; Max size: 30MB per image; Max files: 14
["https://example.com/reference.jpg"]
input.aspect_ratioOptionalstringAspect ratio of the generated image. Defaults to match_input_image when image_input is provided, otherwise 1:1.
Available options:
"16:9"
input.resolutionOptionalstringResolution of the generated image. Higher resolutions produce more detail but take longer to generate. Default: 1K.
Available options:
"1K"
input.output_formatOptionalstringFormat of the output image. Default: jpg.
Available options:
"jpg"
1curl -X POST "https://api.apipass.dev/api/v1/jobs/createTask" \
2 -H "Content-Type: application/json" \
3 -H "Authorization: Bearer YOUR_API_KEY" \
4 -d '{
5 "model": "google/nano-banana-2",
6 "callBackUrl": "https://your-domain.com/api/callback",
7 "input": {
8 "prompt": "A serene alpine lake reflecting snow-capped mountains at golden hour, photorealistic",
9 "aspect_ratio": "16:9",
10 "resolution": "1K",
11 "google_search": false,
12 "image_search": false,
13 "output_format": "jpg"
14 }
15 }'1{
2 "code": 200,
3 "message": "success",
4 "data": {
5 "taskId": "task_12345678"
6 }
7}codeStatus code, 200 for success, others for failure
messageResponse message, error description when failed
data.taskIdTask ID for querying task status and results

ByteDance
ByteDance
Starting from
0 credits
ByteDance
ByteDance
Starting from
0 credits
ByteDance
ByteDance
Turn text, images, video clips, and audio into cinematic, audio-synced videos with a single API call.
Starting from
0 credits
MiniMax
MiniMax
Starting from
0 credits
ByteDance
ByteDance
Generate cinematic 4K AI videos up to 30 seconds with Seedance 2.5's multimodal power.
Starting from
0 credits
Google Veo 3.1 is the latest state-of-the-art generative video model designed to transform text and image prompts into high-fidelity cinematic visuals. Building upon its predecessors, Veo 3.1 features significant upgrades in prompt adherence, visual realism, and native audio generation, creating 8-second clips with synchronized sound effects and dialogue.
Starting from
0 credits
wan
wan
Starting from
0 credits
Runway
Runway
Starting from
0 credits
Kling
Kling
Transform text prompts and images into high-quality videos with advanced AI. Generate cinematic content with customizable duration, aspect ratio, and audio capabilities.
Starting from
0 credits
MiniMax
MiniMax
Generate high-quality AI videos from text prompts or images with MiniMax's Hailuo 2.3 model. Support for both 6s and 10s durations, 768p and 1080p resolutions, with intelligent prompt optimization.
Starting from
0 credits
grok
grok
Grok Imagine is xAI’s AI video generation model for creating high-quality videos with realistic motion, native audio, and synchronized speech. Grok Imagine 1.5 introduces faster generation speeds, improved motion physics, and stronger visual consistency for modern AI video applications.
Starting from
0 credits
Luma
Luma
Starting from
0 credits
Kling
Kling
Kling v3 video generation model, supports text-to-video and image-to-video, up to 15 seconds, 1080p pro mode
Starting from
0 credits
wan
wan
Generate videos with cinematic consistency by reimagining existing footage through high-fidelity visual transformation and structural control.
Starting from
0 credits
grok
grok
Starting from
0 credits
kling
kling
Starting from
0 credits