AI talking avatar: turn a photo and an audio track into a lip-synced video
This tool takes a photo of a face and an audio track and generates a video of that photo talking, with the mouth and expression synced to the audio. On VdoBloom it is the Avatar tab at /dashboard/video-creation/avatar/, and it is not one model but a choice of seven, each with its own inputs, duration limit and price: Kling Avatar Standard, Kling Avatar Pro, InfiniTalk, WAN Speech-to-Video Turbo, Omnihuman 1.5, PixVerse LipSync and Volcengine Video-to-Video Lip Sync.
For five of the seven models, the input is exactly what it sounds like: one photo plus one uploaded audio file. Two models work differently. Volcengine Video-to-Video Lip Sync takes an existing video instead of a photo and re-syncs its lips to a new audio track, useful when you already have footage and just need the mouth movement to match different audio. PixVerse LipSync also takes a video rather than a photo, but it is the one model here that can skip an audio upload entirely — you can type text and pick a TTS voice instead, and the audio is generated for you.
The two Kling Avatar models additionally require a text prompt describing what the avatar should do (for example, "the lady is talking"), on top of the photo and audio — the interface will not let you submit one of those without it. The other five models treat the prompt as optional. Pricing is per second of audio across every model, so cost scales directly with how long your audio clip is, not with the video's resolution alone.
What this tool does
- Seven avatar/lip-sync models are selectable: Kling Avatar Standard, Kling Avatar Pro, InfiniTalk, WAN Speech-to-Video Turbo, Omnihuman 1.5, PixVerse LipSync and Volcengine Video-to-Video Lip Sync.
- Five models take a photo plus an audio file; Volcengine and PixVerse LipSync take an existing video instead of a photo.
- PixVerse LipSync is the only model that can generate the audio itself from typed text and a chosen TTS voice, instead of requiring an uploaded audio file.
- The two Kling Avatar models require a written text prompt describing the avatar's action; the other five models treat the prompt as optional.
- Kling Avatar Standard renders at a fixed 720p and Kling Avatar Pro at a fixed 1080p — resolution is not user-selectable on either.
- Every model bills per second of audio: Kling Standard is about 3.18 credits/second (96 credits for 30 seconds), Kling Pro about 6.37 credits/second (191 credits for 30 seconds).
- Kling Avatar Standard and Pro accept up to 5 minutes (300 seconds) of audio; Omnihuman 1.5 and Volcengine cap at 60 seconds; InfiniTalk and WAN Speech-to-Video Turbo cap at 120 seconds; PixVerse LipSync caps at 300 seconds.
- InfiniTalk and WAN Speech-to-Video Turbo carry a 15-credit minimum charge per job even for very short audio; PixVerse LipSync carries a 5-credit minimum.
- Uploads are capped at 10MB for the avatar photo, 100MB (about 5 minutes) for the audio file, and 500MB for a source video on the video-based models.
- Failed generations are refunded automatically — credits are only kept if the video actually completes.
What it does and what it does not do
It does one job — make a face move and lip-sync to an audio track — across seven different underlying models, so you can trade off quality, price, input type and duration limit. Every model produces a video where the mouth and, to varying degrees, facial expression track the audio you provide. This covers narrated presenter videos, greetings, announcements and short talking-head clips from a single photo.
It does not write a script for you — you supply the audio (or, on PixVerse LipSync only, typed text that VdoBloom converts to speech) — and it does not generate a full scene or background; the person and setting come from your source photo or video. It also is not the same tool as VdoBloom's Spokesperson tab (which builds and saves a reusable synthetic persona by gender, age, ethnicity, style and accent) or the Testimonial tab (a UGC-style testimonial video builder) — those are separate, self-contained tools that happen to sit near this one in the sidebar and may use similar underlying avatar technology, but they are not this generic photo-plus-audio tool.
It also does not accept arbitrarily long audio on every model — the two 60-second-capped models (Omnihuman 1.5, Volcengine) will not process a full 3-minute voiceover; you would need Kling Avatar or PixVerse LipSync for anything past 2 minutes.
How it works: seven models, two input shapes
Kling Avatar Standard and Kling Avatar Pro (Kling AI's Avatar 2.0) take a photo, an audio file and a required text prompt describing the intended action. Standard renders at a fixed 720p, Pro at a fixed 1080p — the resolution is locked to the model, not a setting you choose. Both accept up to 5 minutes of audio, the longest ceiling of any model here, and are priced per second from a 300-second anchor (954 credits for Standard, 1910 for Pro at the full 5 minutes).
Omnihuman 1.5 and InfiniTalk are photo-plus-audio talking-avatar models aimed at realistic lip sync at shorter lengths: Omnihuman caps at 60 seconds of audio, InfiniTalk at 120 seconds, and InfiniTalk's price depends on whether you generate at 720p or 480p. WAN Speech-to-Video Turbo is a third photo-plus-audio option with its own exposed controls (frames, guidance scale, safety checker) for more advanced tuning, also capped at 120 seconds and priced by resolution tier.
The remaining two models take video instead of a photo. Volcengine Video-to-Video Lip Sync re-syncs an EXISTING video's mouth movement to a new audio track — useful for redubbing footage you already have — capped at 60 seconds and offered in a Lite (faster) or Basic (higher quality) mode. PixVerse LipSync also starts from a source video, but is the only model in this group that can generate its own audio: instead of uploading a file, you can type up to 200 characters and pick a TTS voice, and the tool synthesizes the speech before lip-syncing it to the video.
Practical use cases
Presenter and spokesperson-style videos from a single headshot: narrate over audio you recorded or had generated elsewhere, and the photo delivers it on camera. This is the most common path across all five photo-based models.
Redubbing an existing video into new audio without reshooting — record a corrected voiceover and re-sync it to the original footage with Volcengine, keeping the visuals but fixing or replacing what was said.
Quick talking clips with no audio recording step at all, using PixVerse LipSync's type-and-pick-a-voice path, when you don't have a microphone or a script recorded but do have a short line of text and a source video.
Long-form narrated content up to 5 minutes on a single avatar image, which only Kling Avatar Standard or Pro support among these seven models — the others top out at 1-2 minutes.
Limits and pricing
Every model here is billed per second of audio (or, for the video-based models, per second of the target audio track), so the same 30-second clip costs a different amount depending on which model renders it. The table below gives representative costs at three audio lengths across all seven models, using each model's own duration cap and rate.
| Model | Input | Max audio | 10s | 30s | 60s |
|---|---|---|---|---|---|
| Kling Avatar Standard (720p) | Photo + audio + prompt | 300s | 32 cr | 96 cr | 191 cr |
| Kling Avatar Pro (1080p) | Photo + audio + prompt | 300s | 64 cr | 191 cr | 382 cr |
| Omnihuman 1.5 | Photo + audio | 60s | 108 cr | 324 cr | 648 cr |
| Volcengine Video-to-Video | Video + audio | 60s | 32 cr | 96 cr | 192 cr |
| InfiniTalk (720p) | Photo + audio | 120s | 48 cr | 144 cr | 288 cr |
| WAN Speech-to-Video (720p) | Photo + audio | 120s | 96 cr | 288 cr | 576 cr |
| PixVerse LipSync | Video + audio or typed text | 300s | 11 cr | 33 cr | 66 cr |
Tips for good results
Trim your audio to the length you actually need before uploading — since every model bills per second, a 90-second file with 60 seconds of silence still costs for the full 90 seconds.
Use a front-facing, well-lit, single-face photo for the five photo-based models; the lip-sync and expression quality depends on how clearly the mouth and jawline are visible in the source image.
For the two Kling models, write a specific action prompt ("the man is speaking directly to the camera, calm expression") rather than leaving it generic — the prompt is required and directly shapes the output.
If you need more than 60 seconds of audio, skip Omnihuman 1.5 and Volcengine and use Kling Avatar, InfiniTalk, WAN Speech-to-Video Turbo or PixVerse LipSync instead, since those two are hard-capped at 60 seconds.
For a quick, low-cost test before committing to a long audio track, PixVerse LipSync's per-second rate and 5-credit minimum make it the cheapest way to preview lip-sync quality on a short clip.
How it works
1.Pick a model
Kling Avatar Standard/Pro, Omnihuman 1.5, InfiniTalk, WAN Speech-to-Video Turbo, PixVerse LipSync or Volcengine Video-to-Video Lip Sync — each has a different input type, duration cap and price.
2.Upload your photo or video
A face photo (10MB max) for most models, or a source video (500MB max, MP4/MOV/MKV) for Volcengine and PixVerse LipSync.
3.Add audio, or type text for PixVerse
Upload an audio file (100MB max, up to about 5 minutes, MP3/WAV/AAC/OGG) — or on PixVerse LipSync only, type up to 200 characters and pick a TTS voice instead.
4.Write a prompt if using Kling
Required for Kling Avatar Standard/Pro (up to 5000 characters describing the action); optional on the other five models.
5.Generate
Cost is calculated per second of audio at generation time and shown before you confirm; a failed job is refunded automatically.
Frequently asked questions
What does an AI talking avatar tool need as input — image, audio, or text?
It depends on the model. Five of the seven models need a photo plus an uploaded audio file. Two need a video instead of a photo. Only one, PixVerse LipSync, can work from typed text and a chosen TTS voice instead of an uploaded audio file.
Can I make a photo talk without recording audio myself?
Only with PixVerse LipSync, and only if you're starting from a video rather than a still photo — you can type up to 200 characters and pick a voice instead of uploading a recording. Every other model here requires an actual audio file.
How long can the audio be?
It depends on the model: Kling Avatar Standard and Pro accept up to 5 minutes, PixVerse LipSync also up to 5 minutes, InfiniTalk and WAN Speech-to-Video Turbo up to 2 minutes, and Omnihuman 1.5 and Volcengine up to 1 minute.
Do I need to write a prompt?
Only for the two Kling Avatar models — the interface requires a text description of what the avatar should do before you can generate. The other five models accept a prompt optionally but do not require one.
Can I control the video resolution?
Not on Kling Avatar — Standard is fixed at 720p and Pro at 1080p regardless of what you choose. InfiniTalk and WAN Speech-to-Video Turbo do let you pick a resolution tier, which also changes their price.
How much does it cost?
Every model bills per second of audio. For a 30-second clip: Kling Standard is 96 credits, Kling Pro 191, Omnihuman 1.5 324, Volcengine 96, InfiniTalk (720p) 144, WAN Speech-to-Video (720p) 288, and PixVerse LipSync 33. Shorter audio costs less on every model, down to a 15-credit floor on InfiniTalk/WAN or a 5-credit floor on PixVerse.
Can this tool re-sync an existing video to new audio instead of animating a photo?
Yes — that is specifically what Volcengine Video-to-Video Lip Sync does. Upload the existing video plus the new audio track and it re-syncs the mouth movement, keeping the original footage otherwise unchanged.
Is this the same as the Spokesperson or Testimonial tools?
No. Spokesperson lets you build and save a reusable synthetic persona (gender, age, ethnicity, style, accent) for repeated talking videos, and Testimonial is a UGC-style testimonial video builder. Both are separate tabs from this generic photo-and-audio avatar tool, even though they may share underlying technology.
What happens if the generation fails?
Credits are refunded automatically if the job fails to start or fails during processing — you are only charged for a video that actually completes.
When should I use something else instead?
If you want a saved, reusable persona rather than a one-off video from a specific photo, use Spokesperson instead. If you want a UGC-style testimonial format, use Testimonial. If you want gentle animation of an old photo with no lip-sync or talking at all, use Animate Old Photos instead — it is a different pipeline built for memory videos, not speech.
Related tools
Animate old photos: turn a still family photo into a gentle memory video
Upload an old photo. AI repairs scratches and color (7 credits, optional), then animates it into a video from 31 credits at 480p, 4 seconds.
AI voice generator: type text, choose a voice, get an MP3
Type text, pick from 4 TTS engines and up to 30 voices, get an MP3 back. ElevenLabs caps at 5,000 characters; 5-6 credits per 1,000 chars.
Genjutsu AI video: change what is in a clip without reshooting it
Upload a 2-30s clip, label a photo with what it replaces, get the same shot back changed. Faces, motion and camera stay. Seedance 2.5, up to 1080p.
Multi-Model AI Video Generator: 54 Variants, One Login
Generate video with 54 model variants across 14 families — Veo 3.1, Kling 3.0, Seedance 2, Wan 2.7, Runway — on one login and one credit balance.
Ready to try it?
Make a talking avatar video