Guide7 min readSeptember 13, 2026

How to Use Gemini Omni Text-to-Video: Specs, Credit Costs and Prompts

A practical guide to Google Gemini Omni Text-to-Video on VdoBloom: real durations, 720p to 4K resolutions, exact credit costs per tier, prompting structure and a worked example.

Gemini Omni Text-to-Video is Google’s Gemini Omni engine running on VdoBloom, and it turns a written prompt into a 4, 6, 8 or 10 second clip at 720p, 1080p or 4K in either 16:9 or 9:16 — costing 36 credits ($1.80) for a 4-second 1080p render up to 120 credits ($6.00) for a 10-second 4K master. It sits in the Gemini Omni family alongside its image-to-video sibling and the faster Omni Flash 1.1 tier, it carries a 4.7 family rating, and it is one of the few text-to-video models on the platform with a genuine 4K output option.

What Gemini Omni Text-to-Video is actually best at

Three things separate this model from the rest of the text-to-video shelf on VdoBloom.

Finished-resolution delivery. Most text-to-video models top out at 1080p. Gemini Omni renders a true 4K tier, which matters when the clip will land on a television, a trade-show screen, a cinema pre-roll, or inside a timeline that will later be cropped and re-framed. Rendering at 4K and cropping to 9:16 in the edit gives you a sharper vertical than generating vertical at 1080p.

Character consistency. The Gemini Omni pipeline on VdoBloom supports character references: you register a character from reference images (and optionally a voice) and then call it in your prompt, so the same person can appear across several clips of the same sequence. Up to three character references can be attached to a single generation.

Native dialogue. Gemini Omni is a multimodal model, not a silent-video model. Voice configurations carry a name, a voice description and example dialogue, so spoken lines come out of the model itself rather than being dubbed on afterwards.

The real specs

These are read straight from the model configuration that powers the picker, not from a marketing page.

SpecGemini Omni Text-to-Video
FamilyGemini Omni (Google)
CapabilityText-to-video only
Durations4, 6, 8, 10 seconds
Resolutions720p, 1080p, 4K
Aspect ratios16:9, 9:16
AudioNative, via voice configurations
Reference inputsUp to 7 images, up to 3 character references
Content filterSTRICT
Family rating4.7
Workspace tabText-to-Video

Note the short aspect-ratio list. There is no 1:1 and no 4:3 here. The model is tuned for the two formats that actually ship: 16:9 for YouTube and web embeds, 9:16 for Reels, Shorts and TikTok.

Exact credit cost, and the pricing quirk worth knowing

Cost is set by duration and resolution together. One credit is $0.05 on the Lite plan ($15 for 300 credits), so the dollar column below is the honest per-render cost at that rate — and it drops further on annual billing.

Duration720p1080p4K
4 seconds36 credits — $1.8036 credits — $1.8084 credits — $4.20
6 seconds48 credits — $2.4048 credits — $2.4096 credits — $4.80
8 seconds60 credits — $3.0060 credits — $3.00108 credits — $5.40
10 seconds72 credits — $3.6072 credits — $3.60120 credits — $6.00

The quirk: 720p and 1080p bill identically at every duration. There is no cost saving whatsoever in choosing 720p, so unless you specifically want a soft, low-resolution look, always select 1080p. Treat 720p as a legacy option, not a budget one.

The second thing the table shows is that 4K is not a proportional surcharge. The jump is a flat 48 credits at every duration, so 4K at 4 seconds is 2.3Ă— the 1080p price while 4K at 10 seconds is only 1.7Ă—. If you know you want 4K, longer clips give you more seconds per extra credit.

A practical budgeting note: a new account gets 10 free credits with no card, which is not enough for a single Gemini Omni render. Use the free credits to test prompt wording on a cheaper model first, then bring the wording that works over to Gemini Omni for the final render. Credit packs start at $2.49 for 75 credits and never expire, so one pack covers two 4-second 1080p tests with change left over. Full tiers are on the pricing page.

How to prompt Gemini Omni Text-to-Video well

Gemini Omni rewards prompts written like a shot description rather than a wish list. A working structure:

  1. Subject and wardrobe in one clause — who is on screen and what they are wearing.
  2. The single action the clip covers. One action per clip. A 6-second render cannot hold a three-beat sequence; split it into three renders and cut them together.
  3. Camera behaviour — slow dolly in, locked-off wide, handheld follow, 35mm lens. Naming a lens length steadies the framing more reliably than naming a mood.
  4. Light and location — overcast north light, practical neon, golden-hour backlight.
  5. Dialogue in quotes, if you want spoken audio, kept short enough to fit the duration. Roughly 12 to 15 words fit comfortably in 6 seconds of natural speech.

Two habits help specifically on this model: write in the present tense, and state what should not move. Gemini Omni tends to animate everything in frame unless told otherwise, so a line like “the background crowd stays still” buys you a cleaner subject.

Because the content filter is STRICT, this model will reject explicit prompts, gore, and sensitive real-person likenesses. If your concept is swimwear, dance or fitness content that keeps getting refused here, the flexible-tier models on VdoBloom handle it instead — see the flexible-filter generator page. Explicit and illegal content stays blocked everywhere on the platform.

A worked example

Goal: a 6-second vertical hero clip for a coffee brand, delivered at 4K so it can be cropped for both a Reel and a website banner.

Settings: Text-to-Video tab, Gemini Omni Text-to-Video, 6 seconds, 4K, 9:16. Cost: 96 credits, $4.80.

Prompt: A barista in a charcoal apron pulls a double espresso into a white ceramic cup on a walnut counter. Slow 35mm push-in on the cup as the crema forms. Warm morning window light from camera left, deep shadows behind. The barista’s hands move; the rest of the frame stays still.

If the first render is close but the push-in is too fast, change only the camera clause and re-run. Changing one clause at a time is the cheapest way to iterate on a model that costs nearly five dollars a pull.

When to pick a different model

  • You are starting from a photo. Use Gemini Omni Image-to-Video instead — same durations, same resolutions, same prices, but it extends a still you already approved. If the photo shows a real person, you must have that person’s consent before animating them.
  • You want the same family but faster. Omni Flash 1.1 runs the same 4/6/8/10 second grid, adds a 360p tier, and covers both text-to-video and image-to-video under one model id.
  • You need longer than 10 seconds in one render. Kling 3.0 Omni Text-to-Video runs 3 to 15 seconds single-shot at up to 4K, and it accepts 1:1.
  • You want a different cinematic house style with dialogue baked in. Veo 3 Quality is the other strict-filter premium option worth testing against this one.

Start a render from the Text-to-Video workspace, where the exact credit price for your chosen duration and resolution is shown before you commit.

Frequently asked questions

Does Gemini Omni Text-to-Video really output 4K?

Yes. 4K is a selectable resolution at all four durations, priced from 84 credits for 4 seconds to 120 credits for 10 seconds. It is a render tier on the model itself, not an upscale bolted on afterwards.

Why do 720p and 1080p cost the same?

Because the provider prices those two tiers identically and VdoBloom passes the tier through rather than inventing a spread. The practical advice is to always choose 1080p over 720p.

Can it generate speech?

Yes. Gemini Omni supports voice configurations with a voice description and example dialogue, and character references can carry a voice, so spoken lines come from the model rather than from a separate dubbing pass.

What happens if a render fails?

Credits are deducted at submission and refunded automatically if the generation fails, with a guard that prevents a refund when the render actually completed. You are not charged for a failed clip.

Can I download without a watermark?

Yes, on any paid plan. Paid plans download watermark-free and include commercial rights; see the watermark-free generator page for details.

How long can a single clip be?

Ten seconds maximum. For longer sequences, generate several clips with the same character reference and edit them together.

Ready to try it?

Create your first AI video in minutes — no credit card required.

Start Creating Free →