Google's Gemini Omni Flash debuts at #
1# on the Artificial Analysis Text to Video and Image to Video Leaderboards, edging out ByteDance's Seedance 2.0 on both
Gemini Omni Flash is the first model in Google's Gemini Omni family, unveiled at Google I/O in May and opened to developers in public preview on June 30. Google positions Omni as a natively multimodal model that can "create anything from any input", starting with video: it accepts text, images, and video as input, generates clips with native audio, and supports conversational editing, where prompts change a video while preserving the rest of the scene. Gemini Omni Flash generates 3 to 10 second clips at 720p and 24 FPS, in 16:9 or 9:16, with longer durations coming soon.
In the Artificial Analysis Video Arena, Gemini Omni Flash debuts at #
1# on both the Text to Video and Image to Video Leaderboards, narrowly ahead of ByteDance's Seedance 2.0 on each.
Gemini Omni Flash is priced at $0.10 per second of generated video ($6.00 per minute), matching Veo 3.1 Fast. The rate is the same for Text to Video and Image to Video.
It is available now in the Gemini API, Google AI Studio, and the Gemini Enterprise Agent Platform, in the Gemini app and Google Flow for consumers, and at no cost in YouTube Shorts and the YouTube Create app.
Congratulations to
@GoogleDeepMind on the release!
See below for comparisons between Gemini Omni Flash and other leading models in the Artificial Analysis Video Arena 🧵