Kling 3.0 — Multi-shot AI video with your characters and sound.
Kling 3.0 is the video model from Kling AI (Kuaishou), launched in February 2026 and built to tell a scene across several shots. On Deepnia you direct up to 5 shots in one 3–15 second clip, keep up to 3 characters or products identical from shot to shot, and turn on “Generated sound” to add voices, sound effects and ambience.
The video tool opens with Kling 3.0 already selected. The exact cost shows before every creation; you sign up when you generate.
Shot 14 s
Shot 2@element_15 s
Shot 34 s
0 s13 s
Illustration of the shot editor: each shot has its own prompt and length.
At a glance
Maker
Kling AI (Kuaishou)
Length
3 to 15 s (5 s by default)
Aspect ratios
9:16, 1:1, 16:9
Modes
Standard, Pro
Shots
1 to 5 per video
Sound
Voices, sound effects and ambience (“Generated sound”)
What Kling 3.0 does best
Several shots in one video
Split your scene into up to 5 shots, each with its own prompt (500 characters) and a length of 1 to 12 seconds, 15 seconds in total. Kling built version 3.0 to follow these sequences, from shot-reverse-shot to cross-cutting.
The same characters from shot to shot
Create up to 3 persistent elements from 2–4 images or one video, then call them in the prompt with @element_1, @element_2, @element_3. A mascot, a product or a character keeps the same look across the whole video.
Sound generated with the picture
Turn on “Generated sound” to get voices, dialogue, sound effects and ambience in the same generation.
Controlled transitions
In single-shot mode, set the start frame and, if you like, the end frame: Kling 3.0 animates the move from one to the other. Useful for a before/after or a product reveal.
Text that stays readable
Kling highlights how 3.0 preserves text from your images: signs, labels and logos stay readable, which matters for product videos.
How to use Kling 3.0 on Deepnia
Open the video tool
The “Try Kling 3.0” button opens Deepnia’s video tool with the model already selected.
Pick single shot or multi-shot
Turn on “Multi-shot” to open the shot editor, then write a prompt and a length for each shot.
Add your images
Add a start frame (and an end frame in single-shot mode) or your persistent elements, called in the prompt with @element_1.
Set it up and generate
Choose the aspect ratio, the length and the Standard or Pro mode, turn on “Generated sound” to add voices and sound effects, check the cost shown, then create.
Kling 3.0 on Deepnia: what you can set
Inputs
Text, start and end frames, persistent elements (images or video)
Length
3 to 15 s (5 s by default)
Aspect ratios
9:16, 1:1, 16:9; with a start frame, the video keeps its aspect ratio
Modes
Standard, Pro
Multi-shot
1 to 5 shots of 1–12 s, 15 s in total, 500 characters per shot
Persistent elements
Up to 3, each made from 2–4 images or one video
Sound
Voices, dialogue, sound effects and ambience with “Generated sound”
Files
Up to 14 files, including 3 videos; images up to 10 MB, videos up to 100 MB
Prompt
Required in single-shot mode; stay under 2,500 characters, the length Kling recommends
Prompt ideas to try
Copy a prompt, open the tool and adapt it to your project.
Multi-shot ad for a food truck
Multi-shot
3 shots
9:16
Sound on
Shot 1 (4 s): wide shot of a busy food truck at dusk, string lights overhead, customers lining up. Shot 2 (5 s): close-up of loaded tacos served hot, steam rising. Shot 3 (4 s): friends laugh as they share the food, the camera slowly pulls back. Sound: street ambience, upbeat instrumental music, laughter.
@element_1 walks into a bright phone shop, hops onto the counter and proudly presents a new smartphone to the camera, then winks. Bright 3D animation style, gentle orbiting camera. Sound: cheerful sound effects, short upbeat jingle.
Smooth transition from the start frame (client with natural hair in the salon) to the end frame (the same client with long beaded braids). The camera slowly circles her, warm salon lighting. Sound: salon ambience, soft music.
@element_1 walks along a Lagos rooftop at sunset wearing @element_2, the fabric moving in the wind; lateral tracking shot, then a close-up on the embroidery. Sound: city wind, light percussion.
The perfume bottle from the start frame slowly rotates on a marble pedestal, orange blossom petals fall around it, golden reflections on the glass, the label perfectly readable. Cream background, studio lighting.
In a bright coworking space, a young founder turns to her colleague and says: "We launch tomorrow." He smiles and replies: "Then let's make it count." Medium shot, then a close-up on their handshake. Sound: office ambience, soft music.
Write one prompt per shot: subject, action, framing and camera move (wide shot, close-up, tracking…).
Give each shot a realistic length: 3 to 5 seconds is enough for a simple action.
Mention every element you add in the prompt, exactly as the tool names it (@element_1, @element_2, @element_3): the tool checks that each one is called before it starts the creation.
Turn on “Generated sound” and describe the sound you want at the end of the prompt: voices, effects, ambience and the dialogue language.
For dialogue, name the language, such as English or Spanish, and put each line in quotation marks.
Test a short version in Standard mode first, then switch to Pro for the final cut.
Good to know
To keep a character or a product identical, create a persistent element from 2–4 sharp images, or from a 3–8 second video showing a single character, as Kling recommends.
Pick 9:16 for Reels, TikTok and Shorts, 16:9 for YouTube and 1:1 for a square feed; with a start frame, crop it to the right format before adding it.
To tell a story longer than 15 seconds, create several videos with the same elements and join them in your editor.
Use images you have the rights to, especially for an element made from a real person.
Kling 3.0 or another model?
Deepnia’s other video models, to pick the right one for your project.
Model
Length
Resolution
Aspect ratios
Best for
Kling 3.0This model
3–15 s
Set by the mode
9:16 · 1:1 · 16:9
Multi-shot AI video with your characters and sound.
Up to 5 shots on Deepnia, each 1 to 12 seconds long, 15 seconds in total. Each shot has its own prompt of up to 500 characters.
What is a persistent element?
A character or an object you define with 2–4 images or a video, then call in the prompt with @element_1, @element_2, @element_3 to keep it identical from shot to shot. You can use up to 3 per video.
Can the video have sound?
Yes: turn on “Generated sound” and Kling 3.0 creates voices, dialogue, sound effects and ambience together with the picture. Describe the sound you want at the end of the prompt.
How do I make my characters speak?
Turn on “Generated sound”, put each line in quotation marks and say who speaks, in what tone and in which language. Kling documents dialogue in Chinese, English, Japanese, Korean and Spanish, with Chinese dialects and English accents.
How long is a Kling 3.0 video?
From 3 to 15 seconds, 5 seconds by default. In multi-shot mode, the total length is the sum of the shots.
What is the difference between Standard and Pro?
They are Kling 3.0’s two generation modes on Deepnia. According to Kling’s documentation, Standard produces 720p video and Pro 1080p: test your shots in Standard, then switch to Pro for the final cut.
Can I set the first and last frame?
Yes, in single-shot mode: add a start frame, then an end frame if you like; Kling 3.0 animates the move from one to the other. In multi-shot mode, the start frame opens the video.
Kling 3.0 or Kling 2.6: which one should I use?
Kling 3.0 to tell a story over several shots, keep characters identical or pick any length from 3 to 15 seconds. Kling 2.6 for a simple 5 or 10 second clip from text or one image.
Ready to try Kling 3.0?
The video tool opens with Kling 3.0 already selected. The exact cost shows before every creation; you sign up when you generate.