AI Presenter Videos for Online Courses, From One Photo and a Script
Most course creators do not want to set up a camera, lights and a microphone for every lesson update. VdoBloom's Course / Training Presenter turns one photo and a typed lesson script into a talking presenter video: the script is voiced with an AI voice, and the photo is animated to speak it with lip sync, natural head movement and hand gestures.
You choose a framing preset — seated at a desk, a plain studio background, or the setting already in your photo — and the presenter is directed to stay waist-up and face the camera like an instructor. Each video runs up to 5 minutes on the standard and pro quality levels. A longer module script is split at sentence ends into lesson-sized parts, which you generate one after another.
What this tool does
- VdoBloom's Course / Training Presenter takes one photo and a typed script and returns a lip-synced talking presenter video; there is no recording step.
- Framing presets are At a desk (seated, waist-up), Plain studio (clean background, waist-up) and Keep my photo (the photo's own setting).
- Standard (720p) and Pro (1080p) quality each take up to 5 minutes of speech per video; OmniHuman 1.5 takes up to 60 seconds.
- Scripts longer than one video are split at sentence ends into lesson parts, up to about 4,200 characters each.
- The voice-over and the video are billed as one charge; if the video fails, both are refunded automatically.
- Credits are bought as one-time packs that never expire, so no subscription is required.
- The voice is ElevenLabs Multilingual v2 with 21 preset voices.
What it costs, in credits
The price is the voice-over plus the video, shown in the Generate button before you start. The voice costs 5 credits per 1,000 characters of script. The video is priced per second of the recorded voice, which the server measures before charging.
| Quality | Resolution | 10 s of speech | 60 s of speech | Longest video |
|---|---|---|---|---|
| Standard | 720p | 32 credits + voice | 191 credits + voice | 5 minutes |
| Pro | 1080p | 64 credits + voice | 382 credits + voice | 5 minutes |
| OmniHuman 1.5 | 1080p | 108 credits + voice | 648 credits + voice | 60 seconds |
Good fits, and where it is not the right tool
It works well for course intros, lesson walkthroughs, onboarding and compliance modules, weekly updates and FAQ answers — anything where one person speaks to camera. It does not add screen recordings, slides or captions; put the presenter clip into your editor or LMS next to your slides.
- Use a waist-up, front-facing, well-lit photo; the framing preset directs the motion but cannot invent a body that is not in the photo.
- Only upload a photo of yourself or of someone who has given consent.
- For a script in another language, Translate & Dub translates it before voicing it.
How it works
1.Upload a presenter photo
A clear, front-facing, waist-up photo works best. Confirm you have the right to use the face.
2.Paste the lesson script
Type or paste the script. If it is longer than one video takes, it is split into lesson parts at sentence ends.
3.Pick framing, voice and quality
At a desk, Plain studio or Keep my photo; one of 21 voices; Standard 720p, Pro 1080p or OmniHuman 1.5.
4.Generate each part
The button shows the credit total first. Rendering takes a few minutes; the finished MP4 lands in your creations.
Frequently asked questions
Do I need to record my voice?
No. You type the script and VdoBloom voices it with an AI voice, then animates your photo to speak it.
How long can one presenter video be?
Up to 5 minutes of speech on Standard and Pro quality, and up to 60 seconds on OmniHuman 1.5. Longer scripts are split into several lesson videos.
What happens if a video fails?
The voice-over and the video are charged together, and a failed video refunds the whole amount automatically.
Is there a subscription?
No subscription is needed. You can buy one-time credit packs that never expire, and pay only for the videos you make. Monthly plans exist for people who prefer them.
Can I use the videos in my paid course?
Yes. You keep ownership of what you generate, as long as you have the rights to the photo you upload.
Related tools
AI talking avatar: turn a photo and an audio track into a lip-synced video
Upload a photo and audio for a lip-synced talking video. 7 models, from 5 credits for a short clip up to 5 minutes of audio on Kling Avatar.
AI voice generator: type text, choose a voice, get an MP3
Type text, pick from 4 TTS engines and up to 30 voices, get an MP3 back. ElevenLabs caps at 5,000 characters; 5-6 credits per 1,000 chars.
A Pay-As-You-Go Alternative to Synthesia and HeyGen for Training Videos
Make training and course presenter videos from your own photo and a typed script. One-time credits that never expire, no subscription, refunds on failure.
Ready to try it?
Make a presenter video