At Google I/O 2026, Google announced Gemini Omni and the Gemini 3.5 series. Gemini Omni was presented as a model able to produce output in any modality from any input, starting with Gemini Omni Flash for video generation and editing, available in the Gemini app from the day of the announcement, and described as a substantial step forward in world understanding, multimodality, and editing. Gemini 3.5 Flash was the first of a new family combining frontier intelligence with the ability to act, positioned to handle multimodal reasoning quickly and cheaply without always paying for the largest model. The announcement moved past treating text, image, video, and audio as separate products toward a single model indifferent to the form of its input and output.