option
Home
News
Google's Gemini Omni Generates Video from Images, Audio, and Text

Google's Gemini Omni Generates Video from Images, Audio, and Text

May 26, 2026
117

Three years ago, Google introduced Gemini with the aim of developing a multimodal large language model—a unified neural network trained on text, images, audio, and video, capable of generating content across all these formats.

At its Google I/O developer conference today, the company advanced toward this vision with Gemini Omni, a new family of multimodal models. Google CEO Sundar Pichai stated that Omni will empower users to "create anything from any input."

Omni's initial focus is video. Users can now combine images, audio, video, and text. Instead of merely stitching these elements together, Omni intelligently reasons across all modalities to produce a coherent output. This results in high-quality videos that demonstrate an understanding of physics, culture, history, and science.

Omni also enables users to edit photos using simple text commands, eliminating the need for complex software, similar to Google's Nano Banana tool.

Google already offers Veo, a dedicated video model that transforms text and images into videos and allows for directing and customizing avatars. However, Nicole Brichtova, Director of Product Management at Google DeepMind, emphasized that today's release represents more than just a Veo update: "It's the next step in merging Gemini's intelligence with the rendering capabilities of our media models."

During a media briefing on Monday, DeepMind's Chief Technologist Koray Kavukcuoglu provided an example: When prompted with "a claymation explainer of protein folding," Omni quickly generated a stop-motion video with a voiceover explaining, "Proteins begin as chains of amino acids. They fold into structures like alpha helices and flat sections called beta sheets, ultimately forming a precise three-dimensional shape."

The long-term vision for Omni is broader, encompassing capabilities like generating images from audio or audio from video.

"When we first announced Gemini, it was our first natively multimodal AI model," Pichai remarked during the briefing. "We knew training it on a combination of text, code, audio, images, and video would lead to a deeper understanding of the world. With world models, AI is evolving from predicting text to simulating reality. Gemini Omni is the next step in that direction."

As part of this release, users will also be able to create videos featuring their own digital avatars—a feature popularized by OpenAI's now-discontinued Sora app with Cameos. To prevent deepfakes, users must complete a dedicated onboarding process, which involves recording themselves while speaking a series of numbers, according to Brichtova. The avatar is then saved for future use.

Additionally, all videos created with Omni will include Google's SynthID digital watermark, allowing users to verify if content was generated using Gemini products.

The first model in the family is Gemini Omni Flash, launching today on the Gemini app, YouTube Shorts, and the AI creative studio Flow. Flash can render 10-second videos. Brichtova clarified that this duration is not a model limitation but a strategic decision to broaden accessibility, anticipating that most users currently prefer shorter clips. Support for longer videos is planned for the near future.

Google appears to be positioning Omni Flash primarily as a consumer tool. During a call with TechCrunch, Brichtova and DeepMind research engineer Gabe Barth-Maron described avatar use cases as personal, such as creating a video of yourself winning an award or visiting the moon, or removing a bystander from a vacation video background.

Barth-Maron summarized it succinctly: "They're like personalized memes."

"We definitely focused on making this easy for consumers to use," Brichtova said. "Not many video models have successfully crossed over to the mainstream consumer market, so this is our attempt to do that."

This ease of use comes with a caveat: Brichtova and Barth-Maron noted that editing prompts must be highly specific. Otherwise, Omni might over-edit or unintentionally alter elements the user intended to keep—a challenge also faced by Nano Banana users.

Google’s Gemini Omni turns images, audio, and text into video — and that’s just the start

Image Credits:Google

Despite its immediate consumer focus, Omni's potential for enterprise and creative applications is evident. Google will make Omni available via API in the coming weeks. The avatar-generation tool—already available on Shorts—is expected to gain traction among content creators. More broadly, an end-to-end multimodal workflow could revolutionize advertising and filmmaking.

Startup Luma AI is developing a similar agentic tool powered by its own "unified" model, capable of generating an entire ad campaign from a brief and a product image.

"We're actually quite proud of the model's text-rendering capabilities, which are very useful for applications like advertising," Brichtova said. "If you need a product placement or even just a slogan, the accuracy is crucial... We certainly anticipate filmmakers and other creators will adopt this model as well."

More professional use cases may be better served by the upcoming Omni Pro model, designed to deliver superior performance across all Omni tasks. Google has not announced a release date for Pro yet, but Brichtova indicated it will launch when "we achieve a significant leap in capability beyond Flash."

Related article
Google rolls out fake call detection to protect against AI deepfake impersonation scams Google rolls out fake call detection to protect against AI deepfake impersonation scams Google announced on Tuesday that Android is launching fake call detection to protect against AI deepfake impersonation scams. The feature is rolling out globally in Phone by Google to Android 12+ devices this month, starting with Pixel devices.As peo
Frontier AI Labs Refuse to Disclose Containment Strategies for Rogue Models Frontier AI Labs Refuse to Disclose Containment Strategies for Rogue Models Recent research indicates that very few leading AI laboratories have published or demonstrated containment response plans. A containment plan defines the procedures for when an AI system attempts to subvert human control, specifying which access righ
Google debuts audio-powered smart glasses at IO 2026, mirroring Meta’s strategy Google debuts audio-powered smart glasses at IO 2026, mirroring Meta’s strategy Google is making a strategic return to the smart glasses market.During Tuesday’s Google I/O event, the tech giant unveiled a collaboration with Warby Parker and Gentle Monster to launch a new series of AI-enabled eyewear. Designed in partnership with
Related Special Topic Recommendations
Image editing AI Object Removal Editors: Clean Up Portraits, Travel, and Product Shots
AI Object Removal Editors: Clean Up Portraits, Travel, and Product Shots

2026 Latest Best Top-rated AI Object Removal Editors for portraits travel product shots! XIX.AI curates a powerful game-changing collection regularly updated with weekly rankings. These tools offer real-world tests to help you quickly remove unwanted elements, boost content quality, and save tons of time without compromising results. Must-try for anyone aiming to unlock their AI creation edge. Explore now!

10 tools
xix.ai
Text-to-speech Best AI Text to Speech Tools for Online Courses
Best AI Text to Speech Tools for Online Courses

2026 Latest Best Top-rated AI Text to Speech Tools for Online Courses are curated by XIX.AI based on rigorous real-world tests and weekly updated rankings. These powerful tools help creators deliver crystal-clear audio content effortlessly, boosting writing efficiency and streamlining course production. Check out the free vs paid comparison to find your perfect fit. Explore now to unlock your AI edge in online education.

10 tools
xix.ai
writing AI Blog Title Tools for Higher Click Through Rates
AI Blog Title Tools for Higher Click Through Rates

2026 Latest Best Top-Rated AI Blog Title Tools for Higher Click Through Rates! XIX.AI has carefully curated a powerful, game-changing collection of top tools that go through rigorous real-world tests. You’ll find a free vs paid comparison, weekly updated rankings, and detailed insights to help you boost your blog’s traffic efficiently. Must-try options are highlighted to help you unlock your AI edge. Explore now!

10 tools
xix.ai
automation Best AI Task Routing Tools for Support Workflows
Best AI Task Routing Tools for Support Workflows

2026 Latest Best Top-rated AI Task Routing Tools for Support Workflows! XIX.AI has curated a highly powerful game-changing collection of must-try solutions, all undergoing rigorous real-world tests and updated weekly. These tools streamline workflows, boost productivity, and help teams deliver faster, more efficient support. Explore now to discover your perfect tool and unlock your AI edge!

17 tools
xix.ai
Academic Research AI Citation and Paper Summary Tools
AI Citation and Paper Summary Tools

2026 Latest Best Top-Rated AI Citation and Paper Summary Tools Curated by XIX.AI. Get powerful game-changing solutions for quick content creation, improved writing efficiency, and boosting productivity. We offer a free vs paid comparison along with real-world tests and weekly updated rankings to help you find the must-try tool that fits your needs perfectly. Explore now to Unlock your AI edge.

10 tools
xix.ai
Productivity Best AI Productivity Tools for Daily Work
Best AI Productivity Tools for Daily Work

2026 Latest Best Top-Rated AI Productivity Tools for Daily Work! XIX.AI has curated a powerful, game-changing selection based on rigorous weekly updated rankings and real-world tests. You’ll find must-try options that boost writing efficiency, streamline content creation, and help you overcome daily work challenges. Get a free vs paid comparison to find the perfect fit for your needs. Explore now to unlock your AI edge!

9 tools
xix.ai
Comments (0)
0/500
OR