Several images, one video
Add up to 7 reference images: a product, an outfit, a character, a setting. Google built Gemini Omni to combine several sources into one coherent clip; say in the prompt what each image is for.
Video model · Google
Gemini Omni is Google’s video model family, announced in May 2026; its first model, Gemini Omni Flash, brings text and images together in one video, with sound. On Deepnia it is the default video model: 4, 6, 8 or 10 s clips in 720p, 1080p or 4K, 9:16 or 16:9, with speech, music and sound effects, and up to 7 reference images to guide a product, an outfit or a setting.
The video tool opens with Gemini Omni already selected. The exact cost shows before every creation; you sign up when you generate.
Add up to 7 reference images: a product, an outfit, a character, a setting. Google built Gemini Omni to combine several sources into one coherent clip; say in the prompt what each image is for.
Pick 720p, 1080p or 4K to suit the use, 720p by default. Google states that 1080p and 4K are produced by upscaling the generated video.
Google says Gemini Omni renders requested text correctly, such as a label or word-by-word captions, in sync with the action. Keep that text short.
Google highlights the model’s grasp of physical forces (gravity, momentum, fluids) and of real-world knowledge, from history to science: useful for explainer videos.
According to Google, Gemini Omni Flash chains several shots by default. You can direct the cuts and the timing in the prompt, or ask for one continuous shot.
Every video comes with its own sound: voices, music and sound effects. Write your prompt in English or French and put the lines in quotation marks to make your characters speak.
The “Try Gemini Omni” button opens Deepnia’s video tool with the model already selected; it is also the default video model.
Attach up to 7 photos from your device or your Deepnia library (10 MB per image). Without images, the video starts from text alone.
State what each image is for, the camera movement and, if needed, the timing (“[0-3s] …”).
Choose the length, resolution and aspect ratio, check the cost shown, then create.
Copy a prompt, open the tool and adapt it to your project.
One continuous shot, no cuts. Close-up of the bottle from the attached image on a wooden counter, condensation droplets on the glass, golden late-afternoon light. The camera slowly pulls back to reveal a busy café in the background. Realistic commercial look, no added text.
Use the attached images as outfit references, not as first frames. Three models walk one after another across a terrace at sunset; each wears one of the reference outfits, with identical patterns and colors. Shoulder-height camera following the walk, fashion magazine style.
Open-air concert at night, facing the sea. The crowd raises their arms under purple and orange lights. In the last two seconds, large white text appears in the center of the frame: "SATURDAY 8 PM". Only one text on screen, clearly readable.
Educational animation in colorful 3D illustration: the water cycle above a coastal village. [0-3s] The sun heats the lagoon and light vapor rises. [3-7s] Clouds form above the palm trees. [7-10s] Rain falls on the roofs and flows back into the lagoon. Three short labels appear at the right moment: "evaporation", "condensation", "rain".
Turn this sketch into a realistic sequence, using the drawing only as a guide for the motion and never showing the drawing: a bike courier weaves between cars on a busy avenue, golden evening light, low tracking camera. One continuous shot.
One continuous shot, no cuts: a slow tracking shot between the tables of a rooftop grill restaurant in Kigali, in the evening. Flames on the grill, golden grilled fish, guests chatting under string lights. Warm light, cinematic look, no on-screen text.
Deepnia’s other video models, to pick the right one for your project.
| Model | Length | Resolution | Aspect ratios | Best for |
|---|---|---|---|---|
| Gemini OmniThis model | 4, 6, 8 or 10 s | 720p–4K | 9:16 · 16:9 | Short clips up to 4K, guided by your images. |
| Veo | 8 s | Set by the mode | 9:16 · 16:9 | 8 s clips with sound, in three quality levels. |
| Seedance 2.5 | 4–30 s | 480p–720p | 16:9 · 9:16 · 1:1 · 3:4 · 4:3 · 21:9 | Videos up to 30 s with sound and up to 16 image and video references. |
| Kling 3.0 | 3–15 s | Set by the mode | 9:16 · 1:1 · 16:9 | A 5-shot editor, persistent characters and optional sound. |
| HappyHorse | 3–15 s | 720p–1080p | 9:16 · 1:1 · 16:9 · 4:3 · 3:4 | Editing an existing 3–60 s video. |
| Grok Video | 6–30 s | 480p–720p | 9:16 · 1:1 · 16:9 · 2:3 · 3:2 | 6–30 s videos, also in 1:1, 2:3 or 3:2. |
A Google video model family, announced on May 19, 2026 at Google I/O. Its first model, Gemini Omni Flash, creates video from text, images and video. On Deepnia it is the default video model, used with text and up to 7 images.
4, 6, 8 or 10 seconds, 6 seconds by default. For a longer video, create several clips and join them in your editor.
Yes: Deepnia offers 3 resolutions (720p, 1080p, 4K), 720p by default. Google states that 1080p and 4K are produced by upscaling the generated video.
Yes, up to 7 images from your device or your Deepnia library, 10 MB each. The model decides how to use each image based on your prompt, so spell it out, for example “as an outfit reference”.
Yes. Write your prompt in English or French and put the lines in quotation marks: Gemini Omni makes your characters speak, with music and sound effects.
Yes. Every video is generated with its own sound: speech, music and sound effects, with nothing to switch on. Describe the voices and the ambience you want in the prompt.
Attach the same reference images to each creation and describe the product or character with the same words. One scene per clip gives the most consistent results.
Gemini Omni to choose the length and the resolution up to 4K, or to combine up to 7 images. Veo for 8-second clips in three quality levels.
The video tool opens with Gemini Omni already selected. The exact cost shows before every creation; you sign up when you generate.