08/06/2026
Gemini Omni is Google’s new multimodal AI model unveiled at Google I/O 2026. It combines Gemini’s reasoning abilities with advanced media generation, allowing it to create content from almost any type of input text, images, audio, and video.
Key Features
🎥 Video Generation from Anything
* Turn text, images, audio, or existing videos into new videos.
* Generates videos with synchronized audio.
✏️ Conversational Video Editing
* Edit videos by simply chatting with the AI.
* Change backgrounds, camera angles, lighting, objects, and scenes using natural language prompts.
🧠 Real-World Understanding
* Uses Gemini’s world knowledge to create more coherent and context-aware content.
* Better consistency in characters, objects, and physical interactions.
🎙️ Multimodal Input
* Accepts combinations of:
* Text
* Images
* Audio
* Video
* Produces video output, with broader output types planned in the future.
🔒 AI Watermarking
* AI-generated content includes Google’s SynthID watermark technology to help identify AI-generated media.
📱 Availability
* Rolling out through the Gemini app, Google Flow, YouTube Shorts, and YouTube Create.
* Available initially as Gemini Omni Flash for paid subscribers.