Skip to main content

Google's new video model blends images, audio and text

By Steven Van ·

Gemini Omni Flash is rolling out in the Gemini app, Google Flow and YouTube Shorts, with conversational editing and digital avatars.

Gemini Omni is Google's first video model built to take any combination of inputs, generating video from images, audio, video and text together, grounded in Gemini's world knowledge. Videos can also be edited through conversation: characters stay consistent, physics hold up, and the scene remembers earlier instructions across multiple turns.

  • Reasons about physics (gravity, kinetic energy, fluid dynamics) and connects that to Gemini's knowledge of history, science and culture, so it can build explainers and storytelling, not just realistic-looking scenes
  • Takes reference images, video, text or audio (voice only for now, with other audio types coming later) and blends them into one cohesive output
  • Supports Avatars, letting people create a digital version of themselves, with their own voice, to generate videos that look and sound like them; editing audio or speech in existing video is still being tested
  • Every video carries an imperceptible SynthID watermark, verifiable through the Gemini app, Gemini in Chrome and Google Search

The first model in the family, Gemini Omni Flash, is rolling out now in the Gemini app, Google Flow and YouTube Shorts, with image and audio output modalities planned for later. Full details are in Google's announcement.

Google Gemini
Google Gemini
Google's AI assistant — write, plan, brainstorm, generate images, and analyze files with one of the most powerful multimodal models.
View Google Gemini →

Read the original announcement →

Read Google's new video model blends images, audio and text on Creators Toolbox