Gemini Omni Image-to-Video vs Text-to-Video: Same Price, Different Job
Both Gemini Omni video variants share identical durations, resolutions and credit costs on VdoBloom. Here is the real difference, the exact per-tier pricing, and which one to pick for each kind of shot.
Gemini Omni Image-to-Video and Gemini Omni Text-to-Video are the same engine with different starting points β identical durations (4, 6, 8, 10 seconds), identical quality tiers (720p, 1080p, 4K), identical aspect ratios (16:9 and 9:16) and identical credit prices β so the only real question is whether you already have the frame you want the shot to start from. If you have an approved still, pick Image-to-Video. If the shot exists only as a sentence, pick Text-to-Video. Nothing else separates them on VdoBloom.
That sounds like a non-answer, so this guide does the opposite of fence-sitting: it shows exactly where each variant wins, what each one costs in credits and dollars, and the two situations where neither is the right pick.
What the two variants actually are
Both sit inside the same Gemini Omni family on VdoBloom, which is rated 4.7 and carries the strictest content label in the catalog. Google filters explicit material, graphic violence and sensitive real-person prompts at the provider level, so this family is built for brand, client and commercial work rather than edgier concepts. If a prompt gets refused, that refusal happens upstream β it is not a VdoBloom setting you can toggle.
Text-to-Video takes a written prompt and nothing else. You describe subject, setting, camera and light, and the model invents the frame along with the motion. Image-to-Video takes one starting image plus a prompt, and animates outward from the composition you supply. The prompt in that second case is a motion brief, not a scene description β the scene is already decided by your upload.
Spec and price, side by side
Credit costs below come straight from the VdoBloom credit tables. Both variants map to the same pricing keys, which is why every row is identical. Dollar figures convert at the Lite-plan rate, where $15 buys 300 credits and one credit is $0.05.
| Spec or tier | Image-to-Video | Text-to-Video |
|---|---|---|
| Input required | One image + prompt | Prompt only |
| Durations | 4, 6, 8, 10s | 4, 6, 8, 10s |
| Quality tiers | 720p, 1080p, 4K | 720p, 1080p, 4K |
| Aspect ratios | 16:9, 9:16 | 16:9, 9:16 |
| Content tier | Strictest | Strictest |
| 720p or 1080p β 4s | 36 credits ($1.80) | 36 credits ($1.80) |
| 720p or 1080p β 6s | 48 credits ($2.40) | 48 credits ($2.40) |
| 720p or 1080p β 8s | 60 credits ($3.00) | 60 credits ($3.00) |
| 720p or 1080p β 10s | 72 credits ($3.60) | 72 credits ($3.60) |
| 4K β 4s | 84 credits ($4.20) | 84 credits ($4.20) |
| 4K β 6s | 96 credits ($4.80) | 96 credits ($4.80) |
| 4K β 8s | 108 credits ($5.40) | 108 credits ($5.40) |
| 4K β 10s | 120 credits ($6.00) | 120 credits ($6.00) |
Two things worth reading twice. First, 720p and 1080p cost exactly the same β there is no reason to render a Gemini Omni draft at 720p to save money, because it saves nothing. Draft at 1080p. Second, 4K is not a small step up: at 10 seconds it is 120 credits against 72, a 67% premium. Treat 4K as a delivery decision, not a default.
When Image-to-Video is the clear winner
Pick Image-to-Video whenever the look of frame one has already been approved by someone. Product photography that went through a retoucher, a campaign key visual, a character design, an architectural render, a still your client signed off on β in all of those cases the composition is the expensive part, and re-describing it in words to a text model is a lossy round trip that will not come back the same.
It is also the better economic choice when you are iterating. A still generated in VdoBloom image generation costs a handful of credits; a 10-second 4K video costs 120. Locking the frame cheaply in images and only then animating it is the single biggest cost saving available in this family.
One rule applies without exception: if the photo you upload shows a real person, you must have that personβs consent to animate them. That is a condition of use on VdoBloom, and the strict content filter on this family will refuse likeness-sensitive material regardless.
When Text-to-Video is the clear winner
Pick Text-to-Video when the shot does not exist yet and camera motion is the point. Anything involving a move through space β a dolly down a corridor, an aerial reveal, a whip pan onto a subject β comes out more convincingly when the model composes the whole sequence rather than extrapolating from a fixed frame. A still constrains the geometry of shot one, and dramatic camera work tends to fight that constraint.
It also wins on speed of exploration. Three written variants cost three prompts. Three image-led variants cost three images plus three renders. For concepting, mood boards and pitch decks, text-first is simply fewer steps. Start in text-to-video and move to image-to-video once a direction wins.
How to prompt each one differently
A Text-to-Video prompt needs four blocks: subject, environment, camera and light. Example: a ceramic coffee cup on a walnut counter, steam rising in a slow curl, morning side light through a window at frame left, camera pushing in slowly on a locked horizontal axis, shallow depth of field, 16:9.
An Image-to-Video prompt should delete the first two blocks. Your still already carries the subject and environment, and restating them invites drift. The same shot as a motion brief: steam rises in a slow curl, camera pushes in slowly, everything else holds still, 6 seconds. Shorter prompts hold composition better here β over-describing a frame the model can already see is the most common cause of a face or a product label changing mid-clip.
A worked example, costed
Say you need a 10-second 4K hero clip of a skincare bottle for a product page. Text-first, you would prompt, render at 4K, dislike the bottle shape, re-render, and burn 240 credits ($12.00) across two attempts before the geometry is right.
Image-first, you generate bottle stills for a few credits each, pick the one that matches the real packaging, then run a single 10-second 4K Image-to-Video pass at 120 credits ($6.00). Same deliverable, roughly half the spend, and the product actually looks like the product. That is the pattern to default to for anything where an object has to be accurate.
When to skip both
Two cases. If you need square, 4:3 or 21:9 framing, neither variant offers it β the list is 16:9 and 9:16 only, and cropping a 4K render to square wastes the resolution you paid for. And if your concept sits in swimwear, dance or fitness territory, the strictest content tier on this family will refuse it; the more flexible tiers in the VdoBloom catalog exist for exactly that work. There is also a third sibling worth knowing about: Omni Flash 1.1 handles both text and image starts and adds a 360p option on top of the same 4/6/8/10-second structure.
Credit costs are shown in the interface before any render starts, and new accounts get 10 free credits with no card, so you can test a 4-second draft before committing. Full plan detail is on the pricing page, where paid plans also remove the watermark and add commercial rights.
Frequently asked questions
Is Gemini Omni Image-to-Video more expensive than Text-to-Video?
No. Both map to the same credit table on VdoBloom: 36 credits for 4 seconds at 720p or 1080p, up to 120 credits for 10 seconds at 4K. The input type does not change the price at any tier.
Can I use both in one project?
Yes, and it is the recommended workflow. Explore concepts in Text-to-Video at short durations, then rebuild the winning shot from an approved still in Image-to-Video for the final render.
Why does 720p cost the same as 1080p?
The underlying credit tiers group them together β only 4K carries a premium. Since the price is identical, render your drafts at 1080p rather than 720p.
Does Image-to-Video keep my original composition exactly?
It animates outward from your frame and holds composition well, but it is a generative model, not a compositor. Short motion-only prompts preserve the frame best; long scene re-descriptions cause drift.
What content does this family refuse?
It carries VdoBloomβs strictest label. Explicit material, graphic violence and sensitive real-person content are filtered by Google before generation, on uploaded images as well as prompts.
Ready to try it?
Create your first AI video in minutes β no credit card required.
Start Creating Free β