Guide7 min readSeptember 13, 2026

AI Voice and Lip-Sync on VdoBloom: ElevenLabs TTS to Kling Avatar 2.0, Step by Step

Generate a voice track at 5-6 credits per 1,000 characters, then drive a face with it. Real per-second avatar pricing for Kling Avatar 2.0, Infinitalk, Omnihuman 1.5, Wan and Video Lip Sync, plus a worked one-minute example.

The voice and lip-sync workflow on VdoBloom is two separate jobs: generate the voice track on the Audio tab with ElevenLabs, Gemini 3.1 Flash TTS or xAI TTS at 5–6 credits per 1,000 characters, then drive a face with that audio on the Avatar tab, where Kling Avatar 2.0 Standard costs about 3.18 credits per second of audio and Infinitalk at 480p costs 1.2 credits per second. A 60-second voiced avatar clip therefore lands anywhere between roughly 77 credits and 653 credits depending on which avatar model you pick β€” a spread of more than 8x that is worth understanding before you build a whole series.

If you animate a photograph of a real person, you must have that person’s consent before you generate, publish or advertise with the result.

Step 1: make the voice

Voice generation lives on the text-to-speech tab. Four models are exposed, and all of them are priced per 1,000 characters of input text, rounded up β€” a 1,200-character script bills as two units, and a 40-character line still bills as one.

Voice modelCredits per 1,000 charactersCost per 1,000 charsBest for
ElevenLabs Multilingual V25$0.25The default for narration and non-English scripts
ElevenLabs Turbo 2.55$0.25Same price, faster return β€” good for iterating on a read
Gemini 3.1 Flash TTS6$0.3030 voices across 70+ languages; the widest voice choice
xAI TTS6$0.306 expressive voices; character and delivery over neutrality

For scale: spoken English runs at roughly 900 to 1,000 characters per minute of finished audio. So a one-minute voiceover is one billing unit β€” 5 credits, about $0.25, on ElevenLabs. Voice generation is the cheap half of this workflow by an order of magnitude. Iterate freely on the script; it is the avatar step that costs money.

Three habits materially improve the read. Write the punctuation you actually want heard, because commas and full stops set the pacing. Spell out numbers, dates and acronyms the way you want them spoken. And break a long script into paragraph-sized generations, so one bad sentence does not force you to re-render four minutes of audio.

Step 2: drive the face

The Avatar tab takes an image plus an audio file and returns a lip-synced video. You upload the avatar image, upload the audio you just generated (or any audio you own), pick a model, and generate. One option in the list is different: Video Lip Sync takes an existing video plus new audio, rather than a still image, and re-syncs the mouth in footage you already have.

Pricing is per second of audio, and the model you choose is the single biggest cost decision in the whole workflow.

Avatar modelInputMax audioCredits per second30 seconds60 seconds
Infinitalk (480p)Image + audio120s1.236 cr / $1.8072 cr / $3.60
Kling Avatar 2.0 Standard (720p)Image + audio300s3.1896 cr / $4.80191 cr / $9.55
Video Lip Sync (Volcengine)Video + audio60s3.296 cr / $4.80192 cr / $9.60
Infinitalk (720p)Image + audio120s4.8144 cr / $7.20288 cr / $14.40
Kling Avatar 2.0 Pro (1080p)Image + audio300s6.37191 cr / $9.55382 cr / $19.10
Wan Speech to Video Turbo (720p)Image + audio120s9.6288 cr / $14.40576 cr / $28.80
Omnihuman 1.5Image + audio60s10.8324 cr / $16.20648 cr / $32.40

Infinitalk and Wan carry a 15-credit minimum per job, so very short lines never bill below that. Kling Avatar 2.0 is the only option that accepts a five-minute audio file in one pass, which makes it the choice for long-form talking-head content; everything else caps at one or two minutes and forces you to split the script.

A worked one-minute example

Say you want a 60-second spokesperson clip from a single portrait, with a script of about 950 characters.

  • Voice: ElevenLabs Multilingual V2, one billing unit β€” 5 credits ($0.25).
  • Draft pass: Infinitalk at 480p, 60 seconds β€” 72 credits ($3.60). Use it to check timing, framing and whether the read lands, before you spend on quality.
  • Final pass: Kling Avatar 2.0 Standard at 720p, 60 seconds β€” 191 credits ($9.55).

Total for the draft-then-final route: 268 credits, about $13.40, for a finished minute. Going straight to Omnihuman 1.5 instead costs 653 credits, about $32.65 β€” and if the script timing was wrong, you pay that twice. The 480p draft pass is the single highest-leverage habit in this workflow.

Getting a clean result from the avatar step

The image matters more than the model. Use a front-facing portrait where the mouth is unobstructed and reasonably large in frame; head-and-shoulders framing beats a full-body shot, because the model has more pixels to work with around the mouth. Avoid profiles, heavy shadow across the jaw, hands near the face, sunglasses and motion blur. Neutral, closed-mouth expressions sync more reliably than wide smiles.

The audio matters second. Clean, single-speaker audio with no background music syncs best; if you want music under the voice, add it in your editor after the lip-sync, not before. Trim leading silence, because you are billed for it.

If you need a portrait to start from, generate one on the image generator or fix an existing one on the image editor first. That is far cheaper than discovering the source photo was the problem after three avatar renders.

The other direction: audio back to text

The transcription tab runs ElevenLabs speech-to-text at 6 credits per minute of audio, rounded up to the whole minute, so $0.30 a minute. It is most useful at the end of the workflow rather than the start: transcribe the finished voice track to generate caption text, or to confirm the model pronounced everything the way you intended before you commit to the avatar render.

Frequently asked questions

Can I use my own voice recording instead of generating one?

Yes. The Avatar tab takes any audio file you upload, so a recording from your phone or a studio track works exactly the same way and skips the text-to-speech cost entirely. Billing is still per second of that audio.

Which avatar model should I default to?

Infinitalk at 480p for drafts and for anything that will only ever be watched on a phone in a feed. Kling Avatar 2.0 Standard for finished 720p work and for anything over two minutes. Kling Avatar 2.0 Pro when you need 1080p delivery. Omnihuman 1.5 is the premium option for short hero clips where the face fills the shot.

How long can a single avatar clip be?

Five minutes on Kling Avatar 2.0, two minutes on Infinitalk and Wan Speech to Video Turbo, and one minute on Omnihuman 1.5 and Video Lip Sync. For a longer piece, split the script at natural paragraph breaks, render each segment, and join them in your editor.

Can I lip-sync a video I already made?

Yes β€” that is what the Video Lip Sync option is for. It takes an existing video plus a new audio track and re-syncs the mouth, which is how you localise a clip into another language without re-shooting or re-generating it.

Do avatar videos download without a watermark?

On any paid plan, yes, with commercial rights. Plans start at $15 a month for 300 credits, annual billing works out cheaper, and one-time credit packs start at $2.49 for 75 credits and never expire. New accounts get 10 free credits with no card β€” enough for a voice track and a short 480p test, not a finished minute. Full rates are on the pricing page.

Ready to try it?

Create your first AI video in minutes β€” no credit card required.

Start Creating Free β†’