Sound generated with the picture
Google says Veo 3.1 always generates its audio: dialogue, sound effects and ambience come from the same generation as the picture. Put spoken lines in quotation marks and describe the sounds in the prompt.
Video model · Google DeepMind
Veo 3.1 is Google DeepMind’s video model family, launched in October 2025 and joined by Veo 3.1 Lite in March 2026. On Deepnia, Veo creates 8-second clips whose picture and sound (dialogue, sound effects, ambience) are generated together, from text or a photo. Three levels: Lite to test an idea, Fast to produce quickly, Quality for the most polished version.
The video tool opens with Veo already selected. The exact cost shows before every creation; you sign up when you generate.
Google says Veo 3.1 always generates its audio: dialogue, sound effects and ambience come from the same generation as the picture. Put spoken lines in quotation marks and describe the sounds in the prompt.
Lite, Fast and Quality match Veo 3.1 Lite, Veo 3.1 Fast and Veo 3.1. According to Google, Veo 3.1 aims for the highest visual fidelity, Fast for quicker production and Lite for high-volume use at the same speed as Fast. On Deepnia, all three create 8-second clips in 9:16 or 16:9.
Google DeepMind highlights Veo 3.1’s believable physics, realism and adherence to the prompt. Film terms (aerial shot, dolly, close-up) go straight into the text.
Add an image and describe the motion: Veo builds the video from it. Google says version 3.1 follows the prompt more closely and improves picture and sound in image-to-video.
Google added native vertical output in January 2026. Pick 9:16 for Reels, Shorts and TikTok, and 16:9 for YouTube or a website.
The “Try Veo” button opens Deepnia’s video tool with the model already selected.
Under “Mode”, pick Lite to test an idea, Fast to produce quickly and Quality for the final cut.
Stay with text only or add an image to animate (up to 10 MB).
Choose 9:16 or 16:9, describe the picture and the sounds, check the cost shown, then create. The length is always 8 s.
Copy a prompt, open the tool and adapt it to your project.
An eight-second cinematic ad: early morning, a barista pours milk into a cappuccino, the foam draws a leaf, steam rises through a ray of sunlight. The camera slowly pushes in toward the cup. Sounds: hissing espresso machine, clinking cups, soft jazz. No dialogue, no on-screen text.
Animate this photo: the juice bottle stays in the center, droplets slide down the glass, lime slices fall in slow motion around it and the light gently brightens. Sounds: fizzing, clinking ice. No added text.
High-angle view of a lit rooftop at night: the crowd cheers, a big screen lights up behind the stage, golden confetti falls. The camera swoops down toward the stage. Sounds: electronic bass, cheering. No on-screen text.
In a bright gym, a coach finishes a set of squats, straightens up, looks into the camera and says: "Five more minutes. You've got this." Medium shot, light handheld camera. Sounds: upbeat workout music, breathing, gym echo.
Slow aerial shot over a new villa with a pool in Accra, late morning; the palm trees sway slightly and the camera gently descends toward the main entrance. High-end real estate video style. Sounds: birds, light wind, lapping water.
Starting from this photo of the decorated hall, the camera glides slowly between the flower-dressed tables toward the couple’s table, candles flicker, petals drift down. Sounds: soft piano, murmuring guests.
Deepnia’s other video models, to pick the right one for your project.
| Model | Length | Resolution | Aspect ratios | Best for |
|---|---|---|---|---|
| VeoThis model | 8 s | Set by the mode | 9:16 · 16:9 | 8-second clips with sound, in three levels. |
| Gemini Omni | 4, 6, 8 or 10 s | 720p–4K | 9:16 · 16:9 | Choosing the length (4–10 s) and resolution up to 4K, with 7 reference images. |
| Seedance 2.5 | 4–30 s | 480p–720p | 16:9 · 9:16 · 1:1 · 3:4 · 4:3 · 21:9 | Videos up to 30 s with sound and up to 16 references. |
| Seedance 2.0 | 4–15 s | 480p–1080p | 16:9 · 9:16 · 1:1 · 3:4 · 4:3 · 21:9 | 1080p with sound always on, plus image and video references. |
| Kling 3.0 | 3–15 s | Set by the mode | 9:16 · 1:1 · 16:9 | A 5-shot editor, persistent characters and optional sound. |
| Grok Video | 6–30 s | 480p–720p | 9:16 · 1:1 · 16:9 · 2:3 · 3:2 | 6–30 s videos, set to the second, from text or one image. |
Google’s Veo 3.1 family, in three levels: Lite (Veo 3.1 Lite), Fast (Veo 3.1 Fast) and Quality (Veo 3.1). “Quality” is Deepnia’s name for Google’s standard Veo 3.1 model.
Always 8 seconds, whatever the level. For a longer video, create several clips and join them in your editor.
Yes, always: dialogue, sound effects and ambience are generated with the picture, with nothing to switch on. Describe the voices and sounds you want to hear in the prompt.
Yes: add an image (up to 10 MB) and describe the motion you want, and Veo builds the video from it. Say what should move and what should stay still, such as the product or the face.
According to Google, Veo 3.1 (Quality on Deepnia) aims for the highest visual fidelity, Veo 3.1 Fast for quicker generation in everyday production, and Veo 3.1 Lite for high-volume use at the same speed as Fast. On Deepnia, all three create 8-second clips in 9:16 or 16:9.
Put each line in quotation marks, say who speaks and in what tone, then describe sound effects and ambience separately: that is the structure Google recommends. In 8 seconds, one or two short lines are enough.
Veo for an 8-second clip with sound always on, in three levels. Gemini Omni to choose the length (4, 6, 8, 10 s) and the resolution up to 4K, or to combine up to 7 reference images.
Yes. Google says every Veo video carries an invisible digital watermark, SynthID.
The video tool opens with Veo already selected. The exact cost shows before every creation; you sign up when you generate.