On May 19, 2026, at Google I/O 2026, Google officially unveiled Gemini Omni — a brand-new family of multimodal generative AI models that the company is positioning as its boldest creative AI move to date. Gemini Omni is a new family of generative models that accepts text, images, audio, and video as inputs and outputs physics-aware, conversationally editable video, and the first model in the family, Gemini Omni Flash, is live today.
Google's pitch is almost as catchy as the model itself: this is the video version of Nano Banana. Last year, Nano Banana brought Gemini's intelligence to image generation and editing, and since then, it's helped millions of people restore old photos, design from sketches and visualize ideas in ways that weren't possible before. Now Google is replicating that same "magic prompt → magic output" experience for moving pictures. From the start Google built Gemini to be natively multimodal from the ground up, and now they're taking the next step by introducing Gemini Omni, where Gemini's ability to reason meets the ability to create. Omni is the new model that can create anything from any input — starting with video. With Omni, you can combine images, audio, video and text as input and generate high-quality videos grounded in Gemini's real-world knowledge, and you can also easily edit your videos through conversation.
The implications are huge. For the first time, anyone — not just trained editors or VFX artists — can reshape a video clip simply by chatting with an AI. Google just shipped a video AI that lets you edit clips by talking to them — no timeline, no software, no experience needed. This is a serious challenge to OpenAI's Sora, Adobe Firefly, and even traditional NLEs (non-linear editors) like Premiere and CapCut.
What Is Gemini Omni?
Gemini Omni is Google's next-generation multimodal video creation and editing model. Gemini Omni is Google's new AI model that generates and edits video from any mix of text, images, audio, or existing footage — all through natural conversation.
The first publicly available variant is Gemini Omni Flash. Today, Google is rolling out the first model in the Omni family — Gemini Omni Flash — to the Gemini app, Google Flow and YouTube Shorts, and in time, they will support output modalities like image and audio. At Google I/O, the company took a concrete step toward that goal with Gemini Omni, a new family of multimodal models that Google CEO Sundar Pichai says will be able to "create anything from any input." Users can now combine images, audio, video, and text, and rather than simply stitching those inputs together, Omni reasons across all of them to produce a consistent output, resulting in high-quality videos that reflect an understanding of physics, culture, history, and science.
At launch, Flash renders short clips — the first model in the family is Gemini Omni Flash, which will roll out today to the Gemini app, YouTube Shorts, and AI creative studio Flow, and Flash will be capable of rendering 10 seconds of video, which isn't a model limitation, but rather a decision based both on a desire to get it into more hands and an anticipation that most users won't want to make much longer videos yet — though longer video durations are in the pipeline for the near future.
Google has also teased a more powerful tier. The more professional use cases might be better served by the Omni Pro model, which should perform better across all Omni tasks, though Google hasn't said when it will release Pro yet.
Why Did Google Build Gemini Omni? The Core Competitive Edge
From a strategic standpoint, Google sees Gemini Omni as the new flagship for video generation inside Gemini, succeeding Veo inside its consumer surfaces and pushing AI tools to be more genuinely useful and creative for everyday users.
From a product standpoint, three pillars define why Omni matters:
1. Natural-language, multi-turn video editing. Gemini Omni gives you an easier way to edit video — with natural language. Every instruction builds on the last. Your characters stay consistent, the physics hold up and the scene remembers what came before.
In Google's demo video, a violin performance clip was edited across multiple conversational turns: first, the background was changed from a concert hall to a meadow; then, the violin was made invisible while the meadow background was preserved; finally, the camera angle was rotated from the front to the back of the performer — all while the meadow setting and the invisible-violin edit remained intact across every turn.

2. Real-world knowledge. Google stressed that Omni is grounded in Gemini's real-world knowledge, allowing it to reason about history, science, and cultural context rather than simply producing visually convincing but meaningless footage. The model is also notably better at understanding real-world physics — things like gravity, motion, and how liquids behave — which means AI-generated scenes should look significantly less "floaty" than what earlier tools produced. This unlocks a whole new genre: AI-generated explainers and educational videos.
3. Multi-modal references in one shot. Per Google's official launch blog, you can combine an image, a video, and a soundtrack as references for a single generation.
In Google's demo video, a still photo of a space station, a clip of blinking red lights, and a music track with a strong drumbeat were used together as references — and Omni produced a single output in which the circular lights on the space station pulsed in sync with the drumbeat, demonstrating how the model can reason across all three modalities at once.

Features of Gemini Omni
Integrated Directly Into Gemini
True to its "video Nano Banana" billing, Gemini Omni isn't a separate product — it lives inside the Gemini app you already use. Google is launching the first model in the Omni family — Gemini Omni Flash. Gemini Omni Flash is rolling out today to all Google AI Plus, Pro and Ultra subscribers globally through the Gemini app and Google Flow, and it's also rolling out at no cost to users on YouTube Shorts and YouTube Create App starting this week.
Multimodal Reference Inputs
This is arguably Omni's biggest leap. It accepts any combination of inputs — text, images, audio, and existing video — and lets you build or modify a scene through back-and-forth conversation, with each instruction building on the last, keeping characters, lighting, and objects consistent across edits.
You can even upload a style sheet or art-direction reference — a hand-drawn concept showing what certain VFX should look like, with annotations explaining when each effect triggers — and Omni will read both the visuals and the written rules off the sheet, then apply them to a real-world video.
In Google's demo video, a live-action skateboarding clip was uploaded alongside a hand-drawn FX style sheet that specified the visual look of each effect and the physical conditions under which each effect should trigger; Omni interpreted both the artistic style of the sheet and the logical relationship between each drawn effect and the corresponding real-world motion in the skateboarding footage, then re-rendered the clip with the hand-drawn FX correctly overlaid onto each trick at the right moment.

Real-World Knowledge and Physics
Google frames Omni as a model that can use Gemini's world knowledge for historical, scientific, cultural, physical, and narrative context, which makes it interesting for explainers and social education videos, not just visual effects demos. Whether you're visualizing how a volcano forms or explaining Bernoulli's principle, Omni can turn a sentence-long prompt into a coherent, physics-aware visual explanation.
True Video Editing (Not Just Re-Generation)
This is where Omni separates itself from nearly every competitor. Beyond creating videos from scratch, Gemini Omni introduces highly capable video editing driven entirely by natural language. Instead of using traditional, complex video editing software, users can edit videos across multiple conversational turns while the AI maintains character consistency, scene continuity, and physics. According to Google, users can completely reimagine a scene's action, swap out backgrounds, modify camera angles, or alter specific details without losing the thread of the original footage.
In other words, this isn't a model that "loosely references" your input and hallucinates a new video — it actually edits your footage, preserving the camera path, blocking, and motion you originally shot. That effectively absorbs the use cases of dedicated motion-control tools, relighting tools, background-replacement tools, and even stabilization filters.
A caveat from Google: editing prompts currently need to be highly specific, because if a prompt is too vague, the model could over-edit or unintentionally alter parts of the video the user wanted to preserve.
AI Avatars Straight From Your Camera
The rollout includes a personalized digital avatar feature that allows users to generate video likenesses of themselves that look and sound like them. No need to re-upload dozens of reference photos every time — and to prevent abuse, the personal avatar feature lets you create a video clone of yourself, but requires recording yourself reading numbers aloud first — Google's built-in friction against deepfakes.
Gemini Omni vs. Veo 3.1: What's the Difference?
This is the question every creator and developer is asking. Here's the honest breakdown:
| Dimension | Gemini Omni Flash | Veo 3.1 |
|---|---|---|
| Architecture | Unified multimodal model | Specialized video diffusion transformer |
| Inputs | Text + image + audio + video together | Primarily text + image |
| Conversational editing | ✅ Multi-turn, state-preserving | ❌ Re-prompt required |
| Native audio | ✅ Output | ✅ Output (stronger lip-sync) |
| Max resolution | 1080p (Flash) | Up to 4K |
| Max clip length | 10s at launch | Extendable to 1+ minute |
| API availability | Coming "in the coming weeks" | Available now via Gemini API & Vertex AI |
| Best use case | Iterative editing, multi-input creative work | Cinematic finals, long shots, production pipelines |
Gemini Omni is not simply "Veo 4 under a new name." Google now presents Gemini Omni and Veo as separate model surfaces: Gemini Omni sits under Gemini, while Veo remains Google's specialized video generation model line. This matters because many creators were watching Omni as a possible Veo rebrand, but the official release points to a more nuanced answer: Omni is a Gemini creative model family that starts with video, while Veo continues as a dedicated video model family.
The architectural split is the key story. Veo 3 is a specialist — a model trained almost exclusively to take text-and-image input and produce video output, with every layer of the architecture tuned for that one task. Gemini Omni is a generalist that handles multiple input and output modalities through a single unified model. Each approach has real tradeoffs. Specialists optimize harder on their narrow task — that's why Veo 3 hits 4K native resolution and Omni Flash sits at 1080p. Generalists handle inputs the specialist can't even accept — Omni takes audio as an input, not just an output, which means you can feed it a voiceover and have it generate matching video.
On the community side, early Reddit threads and YouTube reviewers echo the same pattern: at the Flash tier, independent reviewers consistently place Gemini Omni Flash's raw cinematic quality one tier below Seedance 2.0 and Kling 3.0, with one tester describing it as "solid mid-to-upper tier" with strong prompt adherence but explicitly noting that visual fidelity lags ByteDance's model. Omni's distinct edge is conversational editing and architectural integration with Gemini reasoning — neither Seedance nor Kling offers multi-turn editing that preserves physics and character state.
Bottom line: Veo 3.1 is more specialized in video duration (extendable to over 1 minute), character consistency (Ingredients to Video), and transition frame generation (First and Last Frame). Omni Flash is more comprehensive in cross-modal input, conversational editing, and physical understanding.
Will Gemini Omni Replace Veo 3.1?
In Google's own words, on the official Gemini video generation page:
"We're always looking for ways to make our AI tools more helpful and creative for users. Gemini Omni is our latest video editing and generation model that will replace Veo in the Gemini app."
But read that carefully — "in the Gemini app." For developers and enterprise users, the picture is very different. Veo has only been replaced by Gemini Omni Flash within the Gemini app product interface. Developers can still effectively use Veo 3.1 and Veo 3.1 Fast via Vertex AI, Gemini API, Google AI Studio, and Flow.
If you open the Gemini app to generate a video, starting in May 2026, the backend will no longer call the Veo series by default, but rather Gemini Omni Flash. For the average user, this is a one-way improvement — stronger multimodal reasoning, conversational editing, and physical understanding are all advantages of Omni. For consumers, Veo has simply stepped quietly into the background.
So: Omni replaces Veo for consumers, but Veo 3.1 remains the production-grade API choice for developers — for now.
What Will Gemini Omni Achieve?
1. Compressing the entire AI video pipeline into a single step. Omni's multi-modal reference capability is arguably its most transformative feature. It doesn't just accept image, video, and audio simultaneously — it understands the logical relationships between them, and can even parse a style sheet's annotations to apply rules onto a video. Where previously a small creative team had to chain text-to-image → image-to-video → motion-control → avatar-injection across multiple models, Omni does all of that in one pass. For lean creative shops and indie creators, this is a game-changer.
2. Lowering the barrier for high-quality creative video. With strong natural-language understanding and stateful multi-turn editing, anyone can now produce polished short-form content. The AI Avatar feature also unlocks personal branding for solo creators — making it easier than ever to build a recognizable IP at influencer scale.
3. Reinventing explainer and educational video production. Just as Nano Banana transformed how explainer images get made, Omni's ability to turn complex concepts directly into didactic video clips will reshape science communication, tutorials, and edtech content.
4. Pressuring traditional editing software. Omni's precise, in-place editing will eat a share of the casual editing market and force traditional NLEs to integrate more conversational AI features. Think of it less like a video generator and more like a creative co-pilot.
5. A coming wave of prompt-to-video templates. Expect to see an ecosystem of ready-to-use Omni prompt templates emerge — for memes, ads, product demos, social hooks, you name it.
Pricing of Gemini Omni
Currently, the only way to access Gemini Omni Flash is via a Google AI subscription (or YouTube Shorts for casual users). Here's the breakdown based on Google's official AI subscription announcement and the Google One AI plans page:
| Plan | Price (US) | Gemini Omni Access | Other Highlights |
|---|---|---|---|
| Free / YouTube Shorts | $0 | Limited free access via YouTube Shorts & YouTube Create | No Gemini app access to Omni |
| Google AI Plus | Entry tier (varies by region) | Limited access to Gemini Omni Flash in Gemini app + Google Flow | Expanded access to Gemini 3.1 Pro; 2x usage limits vs Free |
| Google AI Pro | $20/month | Full access to Gemini Omni Flash + 200 Google Flow Credits | Custom tool creation; YouTube Premium Lite included |
| Google AI Ultra | $100/month (new tier) or $200/month (top tier, reduced from $250) | Highest usage limits + 10,000 or 25,000 Flow Credits | 5× higher usage limit vs Pro; Gemini 3.5 Flash integration; priority access to Google Antigravity; 20TB cloud storage; YouTube Premium |
Google launched a $100/month AI Ultra plan, specifically tailored for developers, technical leads, knowledge workers and advanced creators, and they also reduced the monthly price of the top-tier AI Ultra plan from $250 to $200.
With an AI subscription, you get access to the latest models, including Gemini Omni (AI Plus, Pro and Ultra; global): Gemini Omni is the new model that can create anything from any input, starting with video.
How to access Gemini Omni as a regular user: Subscribe to Google AI Plus (for light use), Google AI Pro at $20/month (recommended for most creators), or Google AI Ultra (for power users). Free access is available exclusively via the YouTube Shorts and YouTube Create apps.
How to Access Gemini Omni
You have four entry points:
1. Gemini app (web or mobile)
Using Gemini Omni on the Gemini app nequires a Google AI Plus, Pro, or Ultra subscription. Just open Gemini, attach your references (image / video / audio), and prompt.

2. Google Flow
Google Flow, aka Google Labs Flow, is Google's AI creative studio. Subscribers get Flow Credits for video generation.

3. YouTube Shorts & YouTube Create app
YouTube offers Gemini Omni free access on Youtube Shorts and YouTube Create app, this version is optimized for short-form creators.

4. Developer API
The Gemini Omni API for Develpoers is Not yet available. According to Google's statement, it's coming in the next few weeks.
Paid subscribers: Gemini Omni Flash is rolling out immediately to Google AI Plus, Pro, and Ultra subscribers within the Gemini app and Google Flow. Free users: The model launches at no cost on YouTube Shorts and the YouTube Create app. Developers and enterprise: Access via APIs will open up in the coming weeks.
One important note for all users: All videos created with Omni include an imperceptible SynthID digital watermark, and you can easily verify that videos were generated with Gemini Omni through the Gemini app, Gemini in Chrome and Google Search.
Is the Gemini Omni API Available on ApiPass?
Not yet — but it's on the roadmap.
As of now, Gemini Omni does not have a public API. However, per Google's official Omni launch announcement, in the coming weeks, they'll also be rolling it out to developers and enterprise customers via APIs.
The moment Google opens the Gemini Omni API, we will prioritize bringing it to ApiPass so our developer community can integrate it with one unified endpoint.
In the meantime, if you need production-grade video generation today, you can already access the full Veo 3.1 family on ApiPass:
- Veo 3.1 Fast — Optimized for throughput and quick iteration.
- Veo 3.1 Lite — The most cost-efficient tier, great for bulk generation.
- Veo 3.1 Quality — The premium tier for highest fidelity output.
Stay tuned to our blog — we'll publish a deep-dive the moment Gemini Omni API access lands on ApiPass.
Conclusion
Gemini Omni represents one of the most consequential AI video launches of 2026 — not because it tops every benchmark (it doesn't), but because it fundamentally reshapes the workflow of creating and editing video. By marrying Gemini's multimodal reasoning with native video output, Google has compressed what used to be a five-tool pipeline into a single conversation.
For consumers, Omni quietly replaces Veo inside the Gemini app and delivers an editing experience that finally feels like talking to a creative collaborator. For developers, Veo 3.1 remains the production-stable choice — but the writing is on the wall: conversational, multi-modal, physics-aware video generation is the new baseline, and every other player from Sora to Kling to Runway will have to respond.
At ApiPass, we'll be tracking Gemini Omni's API rollout closely and bringing it to our platform the moment Google opens the gates. Until then, our Veo 3.1 endpoints give you everything you need to start building today.
The AI video race just got a lot more interesting. Welcome to the Omni era.