How to Use Kling V3 Turbo Text-to-Video: Durations, Credits and Prompt Structure
A hands-on guide to Kling V3 Turbo Text-to-Video on VdoBloom: per-second billing from 3 to 15 seconds, 720p/1080p, 1:1, 9:16 and 16:9 framing, a four-part prompt formula and a worked 8-second example.
Kling V3 Turbo Text-to-Video turns a written prompt into a 3 to 15 second clip at 720p or 1080p in 1:1, 9:16 or 16:9, and VdoBloom bills it per second β 7.2 credits per second at 720p and 9 credits per second at 1080p β so a 5-second 720p draft is 36 credits ($1.80) and a 15-second 1080p final is 135 credits ($6.75). No image, no reference, no first frame: you describe the shot and the model builds everything in it.
What this model is for
Kling V3 Turbo Text-to-Video is the speed tier of Kuaishouβs third-generation line. Its real advantage over most text-to-video models is duration granularity. Nearly every competitor locks you to two or three fixed lengths β 5 and 10 seconds, usually β and you pay for footage you trim away. Here you pick any whole number of seconds from 3 to 15, and because billing is per second, a 7-second cutaway costs 7 seconds of credits.
That makes it a strong fit for editing to a locked timeline: a beat that runs 11 seconds gets an 11-second clip, not a 15-second render you cut down. It is also well suited to the concepting phase, where the useful move is generating six different interpretations of one idea cheaply at 720p and only then committing to a 1080p master.
What it is not: it is not an audio model, and it is not a character-consistency tool. Every generation invents its people fresh, so the same prompt run twice gives you two different faces. If a specific face, product or composition has to survive across shots, start from a still and use Kling V3 Turbo Image-to-Video instead. Like the rest of the Kling family it also applies a strict provider content filter, so keep concepts brand-safe.
Specs and exact credit cost
| Duration | 720p credits | 720p cost | 1080p credits | 1080p cost |
|---|---|---|---|---|
| 3 seconds | 22 | $1.10 | 27 | $1.35 |
| 5 seconds | 36 | $1.80 | 45 | $2.25 |
| 7 seconds | 50 | $2.50 | 63 | $3.15 |
| 9 seconds | 65 | $3.25 | 81 | $4.05 |
| 11 seconds | 79 | $3.95 | 99 | $4.95 |
| 13 seconds | 94 | $4.70 | 117 | $5.85 |
| 15 seconds | 108 | $5.40 | 135 | $6.75 |
Aspect ratios are 1:1, 9:16 and 16:9, picked in the settings panel rather than inherited from anything, which is one practical edge the text-to-video model has over its image-to-video sibling. Credits convert at the Lite rate of $15 for 300 credits, so 1 credit is $0.05; annual billing brings Lite to $10/month and Pro to $49/month at the 2,000-credit tier. One-time packs start at $2.49 for 75 credits and never expire β see pricing for the ladder.
Be aware that the 10 free credits on a new account do not reach the 22-credit minimum for a 3-second clip here. Spend those on a cheaper model to get a feel for the platform β the free generator page covers what they do stretch to.
How to prompt it well
Text-to-video prompting fails in a predictable way: people write a caption instead of a shot. A caption names a subject. A shot names a subject, a lens behaviour, a light source and a duration-appropriate amount of action. Build the prompt in four parts.
- Subject and setting. Concrete nouns with one or two adjectives. A weathered fisherman in a yellow slicker on a wet harbour wall β not a man outside.
- Action, sized to the clip. One continuous action for anything under 8 seconds. Two at most for 12 to 15. Asking for a sequence of three events in 5 seconds is the single most common cause of a warped, smeary result.
- Camera. Name the move and the speed: slow dolly in, locked-off wide, handheld tracking shot at walking pace. If you do not name one, you get a near-static frame.
- Light and look. Overcast grey daylight, hard low sun from the left, neon spill on wet asphalt. This does more for the perceived quality of the render than any other clause.
Two things to leave out. Do not write negatives β no blur, no extra fingers β because naming an object tends to summon it. And do not stack five style words; one reference register, such as documentary handheld or glossy commercial, holds better than a pile of them.
A worked example
Target: an 8-second 9:16 opener for a coffee brand.
Weak prompt: coffee shop video, cinematic, 4k, beautiful. That is a caption with adjectives; the output will be a generic, slowly drifting interior.
Strong prompt: a barista in a dark apron pulls an espresso shot on a brass machine in a narrow city cafe; the camera pushes in slowly from waist height to the portafilter; warm morning light rakes in from a window on the right, steam catching the beam; documentary handheld feel.
Run it at 8 seconds, 720p, 9:16 β 58 credits, $2.90. If the push-in is too aggressive, change slowly to very slowly and re-run for another 58. Once locked, re-run at 1080p for 72 credits ($3.60). Two drafts plus a master lands at 188 credits, about $9.40, and the paid-plan download comes back watermark-free with commercial rights.
When to pick a different model
| What you need | Pick instead | 5-second cost |
|---|---|---|
| Dialogue, ambience or any native sound | Kling 3.0 Omni Text-to-Video | 40 credits at 720p |
| Maximum quality, up to a 4K tier | Kling 3.0 (std / pro / 4K) | 28 std, 36 pro, 134 at 4K |
| A specific face, product or composition preserved | Kling V3 Turbo Image-to-Video | 36 credits at 720p |
| Many cheap variations before you commit | Seedance 1.5 Pro | 8 credits at 720p, 4 at 480p |
| Swimwear, dance, fitness or fashion concepts | Flexible-filter models | varies |
The honest positioning: V3 Turbo Text-to-Video is the iteration model. It is not the cheapest on the platform and it is not the highest ceiling in the Kling family. What it gives you is per-second billing across 3 to 15 seconds, three aspect ratios and quick turnaround β which is exactly what you want when the prompt is still moving. Once the prompt stops moving, Kling 3.0 or Omni is usually the better place to spend the final render.
Frequently asked questions
What does a 15-second clip cost on Kling V3 Turbo Text-to-Video?
108 credits at 720p and 135 credits at 1080p, which works out to $5.40 and $6.75 at the Lite rate of $0.05 per credit. Fifteen seconds is the maximum length the model accepts.
Does Kling V3 Turbo Text-to-Video generate audio?
No, it outputs silent video. If you need native sound from a Kling model, Kling 3.0 Omni Text-to-Video generates audio with the video at 40 credits for 5 seconds at 720p.
Can I use my own image with it?
No. This model is text-only and appears solely on the text-to-video tab. To animate a still, switch to the Image to Video tab and pick Kling V3 Turbo Image-to-Video, which uses your upload as the first frame. If that still shows a real, identifiable person, you need their consent before animating it.
Which aspect ratio should I generate in?
Generate natively in the ratio you will publish in β 9:16 for TikTok and Reels, 16:9 for YouTube and web, 1:1 for feed placements. Cropping a 16:9 render down to 9:16 throws away most of the frame and usually cuts the subject badly.
Why did my prompt get rejected?
The Kling family applies strict provider-side content filtering, so explicit, violent or otherwise sensitive prompts are blocked before generation. Explicit and illegal content stays blocked platform-wide on VdoBloom, but brand-safe concepts that a strict filter still refuses β swimwear, dance, fitness, fashion β often run fine on the flexible-filter models instead.
Where do I find the model?
On the Text to Video tab, under the Kling group, listed as V3 Turbo Text-to-Video. The credit cost for your chosen duration and quality is shown before you generate.
Ready to try it?
Create your first AI video in minutes β no credit card required.
Start Creating Free β