Input video
REMIX REALITY
Change the aesthetic, action, or effect - all from your input video
Tighter control, lower cost
Extend scenes for longer storytelling
Direct it from both ends


References drive the shot

Input image
Input video
Recut without reshooting
Physics that holds at 4K
Transform your world
Change the aesthetic, action, or effect based on your input video
How to generate
Generate using the most advanced video model
1Input Image Reference
Upload reference images to guide your vision
2WRITE THE PROMPT
Use natural language to describe desired scenario and sounds
3Generate with Gemini Omni Flash
Click "Generate" button and Receive high-fidelity video in seconds.
CREATE FROM ANYTHING
Create anything from anything from any input - starting with video
Community over 25 MILLION USERS
Join a global creative network where people generate AI images, share ideas, and inspire each other every day.




Trusted by 5.000+ people worldwide
Pick your plan
Get access to more generations and priority access to new features
Still have questions?
We’ve answered the most frequently asked questions
What is Gemini Omni Flash?
Gemini Omni Flash is Google DeepMind's video generation and editing model. It combines Gemini's reasoning and world knowledge with the ability to create and edit video from any combination of image, text, video, and audio inputs.
How does it work?
You give it a clip, an image, a sketch, or just text — and describe what to do in natural language. It generates or edits the video, maintaining scene consistency across multiple turns. Each instruction builds on the previous one, like directing a conversation with an editor.
How is it different from other video models?
Three things: conversational multi-turn editing with scene consistency, native multimodal inputs (image + text + video + audio in one prompt), and grounding in Gemini's world knowledge — physics, history, culture rendered accurately, not just plausibly.
What input and output formats are supported?
Inputs: image, video, audio, text, sketches. Image: PNG, JPG, WebP up to 20 MB. Video: MP4, MOV, WebM up to 60s. Audio: MP3, WAV up to 30s. Output: MP4 (H.264) in 16:9, 9:16, 1:1, or 4:5.
What resolution and duration can it generate?
Native output: 720p at 24fps. Default clip length is 8 seconds per generation, extendable to 60s via continuation. 1080p upscale available as a post-processing step.
Which languages does multilingual speech support?
English, Spanish, Mandarin, Japanese, French, German, Portuguese, Hindi, Korean, Russian, Arabic. Regional accents and dialects supported. Lip-sync re-renders to match the phonemes of the target language.
How many edit turns are supported per session?
Unlimited multi-turn editing within a session. Scene consistency holds across turns. Each turn produces a new generation; previous turns remain in session history and can be branched or restored.
What safety and provenance signals are included?
Every output carries SynthID — an imperceptible digital watermark — and C2PA Content Credentials embedded in the file metadata. Both persist through standard re-encoding and platform uploads.








