Picture and sound in one pass
Kling AI presents version 2.6 as its first model that generates picture and sound at the same time. Turn on “Generated sound” for voices, sound effects and ambience synced with the action.
Video model · Kling AI
Kling 2.6 (Kling VIDEO 2.6) is the video model from Kling AI (Kuaishou), launched in December 2025 as the brand’s first model to generate picture and sound together. On Deepnia it makes 5 or 10 second clips from text or a photo; turn on “Generated sound” to add voices, sound effects and ambience, with characters who speak or sing in English or Chinese.
The video tool opens with Kling 2.6 already selected. The exact cost shows before every creation; you sign up when you generate.
Kling AI presents version 2.6 as its first model that generates picture and sound at the same time. Turn on “Generated sound” for voices, sound effects and ambience synced with the action.
Kling documents speech, dialogue, narration, singing and rap, on top of sound effects and ambience. Characters speak and sing in English or Chinese, so write your lines in one of those two languages.
Add 1 image: it becomes the video’s first frame, and the clip keeps its aspect ratio. A dish, a shop window or a product comes to life, with or without sound.
A choice of 5 or 10 seconds, 3 social formats (9:16, 1:1, 16:9) and one switch for sound: Kling 2.6 gets straight to the point when a short clip is enough.
The “Try Kling 2.6” button opens Deepnia’s video tool with the model already selected.
Add 1 image of up to 10 MB to use as the start frame. Without an image, the video starts from your text.
Describe the scene, the action and the camera move, then the sound you want: ambience, effects and lines of dialogue in quotation marks.
Pick 5 or 10 seconds and the aspect ratio (without an image), turn on “Generated sound” to add sound, check the cost shown, then create.
Copy a prompt, open the tool and adapt it to your project.
Animate this photo: steam rises gently from the dish, the sauce glistens, a hand places a spoon next to the plate, slight push-in. Sound: soft sizzling, restaurant ambience, clinking cutlery.
A smiling young woman holds a bottle of mango juice in front of a colorful fruit stall and sings a short cheerful line: "Fresh mango, sunny day!" Sunny light, gentle circling camera. Sound: upbeat pop singing, market ambience.
In a bright gym, a coach in sportswear faces the camera while her group does squats behind her. [Coach, energetic]: "Ten more seconds, you've got this!" Medium shot, handheld camera. Sound: breathing, footsteps, gym ambience.
From this photo of the shop window, the camera slowly moves toward the mannequins, the fabrics sway slightly, the window lights switch on one by one. Sound: distant street noise, passing footsteps, a soft door chime.
The camera slowly circles the watch on a stone pedestal, soft reflections on the dial, dark background, studio lighting, smooth and steady motion.
An office decorated for the holidays: a string of lights switches on, gold confetti falls on a table covered with wrapped gifts, slow sideways tracking shot. Sound: sleigh bells, distant laughter, rustling gift wrap.
Deepnia’s other video models, to pick the right one for your project.
| Model | Length | Resolution | Aspect ratios | Best for |
|---|---|---|---|---|
| Kling 2.6This model | 5 or 10 s | Automatic | 9:16 · 1:1 · 16:9 | 5 or 10 s clips, with sound when you want it. |
| Kling 3.0 | 3–15 s | Set by the mode | 9:16 · 1:1 · 16:9 | 3–15 s, up to 5 shots, 3 persistent elements and an end frame. |
| Motion Control | Set by the source video | 720p–1080p | From the source | Copying the moves from a real video onto a character image. |
| Seedance 2.0 | 4–15 s | 480p–1080p | 16:9 · 9:16 · 1:1 · 3:4 · 4:3 · 21:9 | Videos with sound up to 1080p, guided by image and video references. |
| Veo | 8 s | Set by the mode | 9:16 · 16:9 | 8 s clips with sound, in 3 quality levels. |
| HappyHorse | 3–15 s | 720p–1080p | 9:16 · 1:1 · 16:9 · 4:3 · 3:4 | Editing an existing video or combining up to 9 images. |
5 or 10 seconds, 5 seconds by default. Kling recommends 10 seconds for dialogue or a song.
Yes: turn on “Generated sound” and Kling 2.6 creates voices, dialogue, singing, sound effects and ambience together with the picture. Describe the sound you want at the end of the prompt.
Turn on “Generated sound”, then write the line in English or Chinese as [Character, emotion]: "line", the format Kling recommends. Choose 10 seconds for dialogue or a song.
Yes: add 1 image (up to 10 MB). It becomes the start frame and the video keeps its aspect ratio.
Without an image: 9:16, 1:1 or 16:9. With an image: the photo’s aspect ratio.
Create several 10-second clips and join them in your editor. To chain up to 5 shots in one 15-second video, use Kling 3.0.
Start from the same product photo for each creation and describe it with the same words. To keep up to 3 characters or products identical within one video, use Kling 3.0’s persistent elements.
Kling 2.6 for a 5 or 10 second clip that is quick to set up, from text or a photo. Kling 3.0, launched in February 2026, for any length from 3 to 15 seconds, up to 5 shots, 3 persistent elements and start and end frames.
The video tool opens with Kling 2.6 already selected. The exact cost shows before every creation; you sign up when you generate.