Text-to-Video vs Image-to-Video on VdoBloom: Which Mode to Pick and When
Text-to-video invents the scene; image-to-video preserves a real face or product. Which VdoBloom models accept which input, what each route costs in credits, and the cheaper hybrid most projects should use.
Pick text-to-video when the subject does not exist yet and you want to invent the whole scene; pick image-to-video when a face, product or outfit already exists in a photo and has to survive the render unchanged. On VdoBloom that decision also narrows which of the 105+ models you can select, because several accept only text, several require an image, and a third group accepts either. For most real projects the cheapest route is neither pure mode but a hybrid: make the still first, then animate it.
What each mode actually controls
Text-to-video hands the model a blank frame. It invents the subject, the wardrobe, the lighting and the camera in one pass. That freedom is the point β and also the problem. Run the same prompt twice and you get two different people. If your video needs a consistent character across four clips, text-to-video alone will fight you.
Image-to-video hands the model frame one. The subject, the clothing, the background and the lighting are already decided; the model only has to invent motion. Identity holds far better, colour grading stays put, and prompts get shorter because you are describing movement rather than a whole world. The trade is that you cannot change what is in the photo. If the shot is wrong, no prompt will fix it β you have to fix the photo.
If the photo you upload shows a real person, you must have that personβs consent before animating it. That applies to every image-to-video model and every effect tab on the platform.
Which VdoBloom models take which input
This is the practical constraint most people hit first. Several strong models are locked to one mode, and the picker on the text-to-video tab hides the ones that need an image, while the image-to-video tab hides the text-only ones.
- Image-only: Kling V2.5 Turbo, Kling V3 Turbo I2V, Kling 3.0 Omni I2V, Kling 2.6, Runway Gen-4 Turbo, Hailuo 2.3 Pro and Standard, MiniMax H3 I2V, Vidu 2.0, Wan 2.5 I2V, Wan 2.6 Flash, Grok Imagine Video 1.5, HappyHorse image and reference modes, SkyReels V4 I2V, Gemini Omni Video I2V.
- Text-only: HappyHorse T2V, HappyHorse 1.1 T2V, SkyReels V4 T2V, Gemini Omni Video T2V, Kling V3 Turbo T2V, Kling 3.0 Omni T2V, MiniMax H3 T2V.
- Either mode: the Seedance 2 family (Mini, Fast, 2, 2.5), Veo 3 Fast and Quality, PixVerse, Vidu Q2 and Q3, Runway Gen-4.5.
That list matters more than any quality ranking. If you have decided you want Kling V2.5 Turbo, you have also decided you are shooting image-to-video, because there is no text path into that model.
What each route costs
Credits are the honest comparison. One credit is $0.05 on the Lite plan ($15 for 300 credits), so the dollar column below is the real per-clip cost at that rate. Annual billing and larger packs push the effective rate lower.
| Model | Input | Tier | Credits | Cost at $0.05/credit |
|---|---|---|---|---|
| Seedance 2 Mini | Text or image | 480p, 5s | 8 | $0.40 |
| Seedance 2 Mini | Text or image | 720p, 5s | 17 | $0.85 |
| Seedance 2 Fast | Text or image | 720p, 5s | 50 | $2.50 |
| Seedance 2 | Text or image | 1080p, 5s | 203 | $10.15 |
| Veo 3 Fast | Text or image | Standard | 40 | $2.00 |
| PixVerse V5 | Text or image | 720p, 5s | 25 | $1.25 |
| Runway Gen-4 Turbo | Image only | 5s | 20 | $1.00 |
| Runway Gen-4 Turbo | Image only | 10s | 40 | $2.00 |
| Kling V2.5 Turbo | Image only | 5s | 40 | $2.00 |
| Hailuo 2.3 Standard | Image only | 768p, 6s | 15 | $0.75 |
| Vidu 2.0 | Image only | 1080p, 8s | 35 | $1.75 |
Notice that the image-only column is not systematically cheaper or dearer. Hailuo 2.3 Standard at 15 credits undercuts almost every text-to-video tier, while Kling V2.5 Turbo at 40 credits for five seconds costs the same as Veo 3 Fast. Mode does not set price; the model does.
The hybrid route most people should use
Generating a still first is usually cheaper than iterating on video. An image on the image generator costs a handful of credits β Qwen 3 text-to-image is 5, Nano Banana is 6, Seedream 4.5 is 7, GPT Image 2 at 1K is 3 β against 20 to 203 credits for a single video attempt. So you iterate where failure is cheap, lock the frame you like, and spend video credits only once.
It also fixes the consistency problem. One approved still, reused as frame one across four image-to-video clips, gives you the same character in all four. Text-to-video cannot promise that. If the still is close but not right, the image editor is a cheaper correction than another video render.
Worked example: ten concepts, two routes
Say you need ten product-teaser concepts and will ship the best three.
Pure text-to-video. Ten attempts on Seedance 2 Fast at 720p and 5 seconds is 10 Γ 50 = 500 credits, or $25.00. Any concept you dislike, you re-render at full video price.
Hybrid. Ten stills with Qwen 3 text-to-image at 5 credits each is 50 credits. Pick three, animate each with Hailuo 2.3 Standard at 768p and 6 seconds, 15 credits each, for 45 credits. Total 95 credits, or $4.75 β and you saw all ten concepts as images before spending a single video credit.
The hybrid route wins here by a factor of five. It stops winning when the motion itself is the idea β a crowd, a vehicle, a complex camera move β because no still can preview that.
When to override the default advice
- Use text-to-video for establishing shots, abstract or stylised sequences, anything where no real subject is involved, and early exploration when you genuinely do not know what the frame should look like.
- Use image-to-video for anything with a recurring person, any real product, brand assets that must match a spec sheet, and every effect on the platform that animates an uploaded photo.
- Use the hybrid for volume work, client options and anything where you expect to throw most attempts away.
Prompting differs too. A text-to-video prompt describes subject, setting, lighting and camera. An image-to-video prompt should describe motion and camera only β re-describing what is already in the photo tends to make the model redraw it, which is exactly the drift you chose image-to-video to avoid. Ready-made examples for both shapes live in the prompt library.
New accounts get 10 free credits with no card, which is enough to run a 480p Seedance 2 Mini draft and see the difference for yourself before you look at plans.
Frequently asked questions
Is image-to-video always more accurate than text-to-video?
For identity and product likeness, yes β the model starts from your pixels. For motion quality the two modes are comparable; that depends on the model you picked, not the input type.
Can I use the same prompt for both modes?
You can, but you should not. Text-to-video prompts describe a whole scene. Image-to-video prompts should describe only movement and camera, because the scene is already fixed by the photo.
Which mode is cheaper on VdoBloom?
Neither. Cost is set by the model, resolution and duration, not the input. Hailuo 2.3 Standard at 768p and 6 seconds is 15 credits with an image; Seedance 2 at 1080p and 5 seconds is 203 credits either way.
Why can I not see Kling V2.5 Turbo on the text-to-video tab?
Because it is image-to-video only. The picker hides models that cannot accept the current tabβs input rather than letting you submit a job that would fail.
Do paid plans remove the watermark on both modes?
Yes. Paid plans download watermark-free with commercial rights, regardless of which input mode produced the clip.
Ready to try it?
Create your first AI video in minutes β no credit card required.
Start Creating Free β