No output yet. Submit the form to generate content.
Complete guide to using Gemini Omni
Turn a reference image and a short description into a reusable AI character that you can plug into every Gemini Omni video generation on ApiPass.

Gemini Omni Character on ApiPass lets you create a reusable character resource from a short description and a reference image. Instead of re-describing a character or re-uploading photos on every call, you register the character once and receive a compact resource you can reuse in any future Gemini Omni Video request — behind the same API key you already use for the rest of the ApiPass model catalog.
Register a character a single time and reference it in every future Gemini Omni Video call — no re-uploading photos, no retyping descriptions.
Anchor a character to a reference image so the same face, style, and identity carry cleanly from clip to clip and scene to scene.
Character resources pair naturally with voice profiles created through Gemini Omni Audio, so a character can carry both a consistent look and a consistent voice across an entire series or campaign.
Each created character returns a compact resource that's easy to persist in your own database and surface in your product's UI.
Character creation lives inside the same ApiPass surface as Gemini Omni Video and Gemini Omni Audio, so you can build end-to-end character-led pipelines without juggling multiple providers.
Register a new character from one clear reference image plus a natural-language description covering appearance, style, and personality.
Give each character a name when creating the resource so it's easy to identify inside your own library or UI.
Pair a character with voice profiles from Gemini Omni Audio so vocal persona is associated with the character from the start.
The API returns a reusable character resource that can be attached to any Gemini Omni Video request to drive character-consistent generations on demand.
Multiple registered characters can be combined in a single Gemini Omni Video request, so scenes with more than one recurring character stay on-model together.
Create your AI avatar once on ApiPass, then let it host your videos for you. Whether you're running a YouTube channel, a TikTok series, or a weekly podcast, your character shows up consistently in every episode — same face, same voice, same vibe. No more setting up cameras, no more bad hair days. Just write your script, and let your AI self do the talking.
A single approved character that appears across ad variants, product tutorials, and localized markets, keeping brand identity intact everywhere it shows up.
Protagonists, sidekicks, and hosts for episodic short-form content and serialized story worlds, all staying on-model from one episode to the next.
Consistent teachers, narrators, and hosts for courses, explainers, and training modules, so every lesson feels like it's delivered by the same familiar face.
Building Episodic Series Where The Same Host Or Protagonist Must Appear Reliably Across Every Video.
Maintaining An On-Brand Spokesperson Or Mascot That Stars In Every Ad, Explainer, And Localized Variant.
Designing A Consistent Instructor Or Presenter Persona That Fronts Entire Course Catalogs.
Registering Recurring Characters That Appear Across Cinematics, Trailers, And In-App Story Sequences.
Offering End-Users A "My Characters" Library To Save, Manage, And Reuse Their Own AI Identities.
Create an ApiPass account and generate an API key from your dashboard to unlock the full Gemini Omni family — character, video, and audio — behind a single credential.
Pick one clear reference image, write a short description covering appearance and style, and choose a character name. Optionally pair the character with a voice profile from Gemini Omni Audio for a fully realized identity.
Call Gemini Omni Character to register the character, then attach it to any future Gemini Omni Video request to generate consistent, character-led videos at scale.
All APIs require authentication via Bearer Token.
Authorization: Bearer
Submit a new Nano Banana 2 image generation or editing task
The API accepts a JSON payload with the following structure:
1{
2 "model": "string",
3 "callBackUrl": "string (optional)",
4 "channel": "auto",
5 "input": {
6 // Input parameters
7 }
8}modelRequiredstringThe model name to use for generation
"google/nano-banana-2"
callBackUrlOptionalstringCallback URL for task completion notifications. If omitted, no callback will be sent.
"https://your-domain.com/api/callback"
channelOptionalstringYou may specify the corresponding provider within APIPASS via the channel parameter; these providers handle the actual image and video generation tasks. APIPASS currently offers three provider options:
The default value for the channel parameter is auto. When enabled, APIPASS automatically allocates tasks across available providers based on real-time pricing and stability metrics to balance minimal cost and reliable performance. Retain the default auto value unless you have custom routing requirements.
Available options:
auto
The input object contains the following parameters:
input.promptRequiredstringA text description of the image you want to generate
Describe the subject, style, lighting, and composition for best results
"A serene alpine lake reflecting snow-capped mountains at golden hour, photorealistic"
input.image_inputOptionalarray(URL)Input images to transform or use as reference. Supports up to 14 images.
Accepted types: image/jpeg, image/png; Max size: 30MB per image; Max files: 14
["https://example.com/reference.jpg"]
input.aspect_ratioOptionalstringAspect ratio of the generated image. Defaults to match_input_image when image_input is provided, otherwise 1:1.
Available options:
"16:9"
input.resolutionOptionalstringResolution of the generated image. Higher resolutions produce more detail but take longer to generate. Default: 1K.
Available options:
"1K"
input.output_formatOptionalstringFormat of the output image. Default: jpg.
Available options:
"jpg"
1curl -X POST "https://api.apipass.dev/api/v1/jobs/createTask" \
2 -H "Content-Type: application/json" \
3 -H "Authorization: Bearer YOUR_API_KEY" \
4 -d '{
5 "model": "google/nano-banana-2",
6 "callBackUrl": "https://your-domain.com/api/callback",
7 "input": {
8 "prompt": "A serene alpine lake reflecting snow-capped mountains at golden hour, photorealistic",
9 "aspect_ratio": "16:9",
10 "resolution": "1K",
11 "google_search": false,
12 "image_search": false,
13 "output_format": "jpg"
14 }
15 }'1{
2 "code": 200,
3 "message": "success",
4 "data": {
5 "taskId": "task_12345678"
6 }
7}codeStatus code, 200 for success, others for failure
messageResponse message, error description when failed
data.taskIdTask ID for querying task status and results

ByteDance
ByteDance
Starting from
0 credits
ByteDance
ByteDance
Starting from
0 credits
ByteDance
ByteDance
Turn text, images, video clips, and audio into cinematic, audio-synced videos with a single API call.
Starting from
0 credits
MiniMax
MiniMax
Starting from
0 credits
ByteDance
ByteDance
Generate cinematic 4K AI videos up to 30 seconds with Seedance 2.5's multimodal power.
Starting from
0 credits
Google Veo 3.1 is the latest state-of-the-art generative video model designed to transform text and image prompts into high-fidelity cinematic visuals. Building upon its predecessors, Veo 3.1 features significant upgrades in prompt adherence, visual realism, and native audio generation, creating 8-second clips with synchronized sound effects and dialogue.
Starting from
0 credits
wan
wan
Starting from
0 credits
Runway
Runway
Starting from
0 credits
Kling
Kling
Transform text prompts and images into high-quality videos with advanced AI. Generate cinematic content with customizable duration, aspect ratio, and audio capabilities.
Starting from
0 credits
MiniMax
MiniMax
Generate high-quality AI videos from text prompts or images with MiniMax's Hailuo 2.3 model. Support for both 6s and 10s durations, 768p and 1080p resolutions, with intelligent prompt optimization.
Starting from
0 credits
grok
grok
Grok Imagine is xAI’s AI video generation model for creating high-quality videos with realistic motion, native audio, and synchronized speech. Grok Imagine 1.5 introduces faster generation speeds, improved motion physics, and stronger visual consistency for modern AI video applications.
Starting from
0 credits
Luma
Luma
Starting from
0 credits
Kling
Kling
Kling v3 video generation model, supports text-to-video and image-to-video, up to 15 seconds, 1080p pro mode
Starting from
0 credits
wan
wan
Generate videos with cinematic consistency by reimagining existing footage through high-fidelity visual transformation and structural control.
Starting from
0 credits
grok
grok
Starting from
0 credits
kling
kling
Starting from
0 credits