Comparison7 min readSeptember 13, 2026

HappyHorse 1.1 Reference-to-Video vs Text-to-Video: Same Price, Different Job

Both HappyHorse 1.1 modes cost 45 credits for 5 seconds at 720p and share all nine aspect ratios. The difference is subject consistency - here is exactly when to upload references and when to prompt from scratch.

HappyHorse 1.1 Reference-to-Video and HappyHorse 1.1 Text-to-Video cost exactly the same on VdoBloom β€” 45 credits for a 5-second 720p clip, 173 credits for a 15-second 1080p one β€” and share every duration, resolution and aspect ratio. The only real difference is the input: reference-to-video takes uploaded images of a subject and keeps that subject recognisable in a brand new prompted scene, while text-to-video builds the whole frame from your words alone. If you need the same face, mascot or product to survive across several clips, pick reference-to-video. If the shot has no fixed subject to protect, text-to-video is one less thing to prepare.

The two modes are the same engine

Both modes sit under the HappyHorse family in the VdoBloom model picker, both are rated 4.8, and both carry the MODERATE content label β€” the standard middle tier, which covers portraits, fashion, dance, fitness and product work without the friction some stricter models apply. Explicit and illegal content stays blocked on every model on the platform.

Because they run on the same HappyHorse 1.1 backbone, the settings panel is identical. Duration is selectable at every whole second from 4 to 15, output is 720p or 1080p, and framing covers nine aspect ratios: 16:9, 9:16, 1:1, 4:3, 3:4, 4:5, 5:4, 21:9 and 9:21. That nine-ratio spread is the headline feature of the 1.1 revision β€” true ultrawide 21:9 and its vertical mirror are formats most generators simply do not offer, and 4:5 and 5:4 are the feed-native crops that stop you letterboxing a social post by hand.

Specs and credit cost, side by side

VdoBloom prices both modes off one shared credit table, so nothing below is an estimate. Credits convert at 1 credit = $0.05 on the Lite plan ($15 for 300 credits); annual billing brings that down, and the Pro 2,000-credit tier at $49 a month annual works out cheaper still. Full plan detail is on the pricing page.

AttributeReference-to-Video 1.1Text-to-Video 1.1
Required inputReference image(s) + promptPrompt only
Durations4–15s, 1-second steps4–15s, 1-second steps
Resolutions720p, 1080p720p, 1080p
Aspect ratios16:9, 9:16, 1:1, 4:3, 3:4, 4:5, 5:4, 21:9, 9:2116:9, 9:16, 1:1, 4:3, 3:4, 4:5, 5:4, 21:9, 9:21
Content tierMODERATEMODERATE
5s at 720p45 credits ($2.25)45 credits ($2.25)
10s at 720p90 credits ($4.50)90 credits ($4.50)
15s at 720p135 credits ($6.75)135 credits ($6.75)
5s at 1080p58 credits ($2.90)58 credits ($2.90)
10s at 1080p116 credits ($5.80)116 credits ($5.80)
15s at 1080p173 credits ($8.65)173 credits ($8.65)

The pattern is linear: 720p bills at 9 credits per second, 1080p at roughly 11.5 credits per second. A second of 1080p costs about 2.5 credits more than a second of 720p β€” twelve and a half cents. That makes iterating at 720p and finishing at 1080p cheap enough to be the default habit rather than a discipline you have to enforce.

What reference-to-video actually does

The common mistake is to assume reference-to-video is image-to-video with a different name. It is not. Image-to-video treats your upload as the literal first frame and moves it. Reference-to-video treats your uploads as a description of a subject β€” this is who the person is, this is what the bottle looks like, this is the style to hold β€” and then generates a completely new scene from your prompt with that subject inside it.

The practical consequence is casting. You can put one character in an ultrawide 21:9 establishing shot and a 9:16 vertical cut from the same reference set, and they will still read as the same character. With text-to-video you would get a plausible person in each shot, but a different plausible person each time.

If you are uploading reference images of a real person, you must have that person’s consent before generating video of them β€” that applies whether the subject is a client, a colleague or a friend.

When text-to-video is the better call

Reference images are overhead. Gathering them, cleaning them up and re-uploading them for every run costs time that is only worth spending when consistency matters. Reach for HappyHorse 1.1 Text-to-Video when:

  • The shot has no recurring subject β€” landscapes, abstract motion, crowds, weather, b-roll texture.
  • You are exploring a concept and want ten different looks, not one repeated look.
  • The subject is generic on purpose: an anonymous hand, a silhouette, a passer-by.
  • You want an ultrawide 21:9 plate and have no source material at all.

Reach for HappyHorse 1.1 Reference-to-Video when the same face, mascot, garment or SKU has to appear in shot two, shot three and shot nine and still look like itself.

A worked example

Say you are cutting a four-clip promo for a skincare bottle. Clip one is a 21:9 ultrawide establishing shot of the bottle on wet stone; clips two to four are 9:16 verticals for Reels.

Doing all four in text-to-video means describing the bottle four times and accepting four slightly different bottles β€” different label, different cap, different proportions. Doing them in reference-to-video means uploading three clean product shots once, then writing four prompts that describe only the scene and the camera.

Cost at 720p, 8 seconds each: 72 credits per clip, 288 credits for the set β€” $14.40 at Lite rates. Upgrade the hero shot alone to 1080p and it becomes 93 credits, taking the set to 309 credits. That is the whole reason to prototype at 720p: you only pay the 1080p premium on the shots that ship.

How to prompt each one well

For text-to-video, front-load the subject, then the action, then the camera, then the light. Something like: a lone red kayak crossing a glass-flat lake at dawn, slow push-in from behind, low mist, warm rim light. Fifteen seconds is a long time for one continuous idea, so give longer clips a beat structure β€” one move, then a second move β€” rather than repeating a single instruction.

For reference-to-video, invert it. The references already carry identity, so do not re-describe the subject in detail; you will only fight the images. Describe the environment, the action and the framing, and refer to the subject in short: the woman from the reference walks into frame from the left, tracking shot, late-afternoon sun. Use two or three references that agree with each other β€” consistent lighting and angle in your uploads produces a more stable subject than one hero shot plus two mismatched ones.

When to pick a different model entirely

If you already have the exact frame you want to move β€” a finished render, a photograph, an image you made in the image generator β€” neither of these is right. You want HappyHorse 1.1 Image-to-Video, which animates the upload itself and prices from the same shared table at 45 credits for 5 seconds at 720p. Note that it does not expose an aspect-ratio control, because the source image sets the frame.

Both modes in this comparison live in the VdoBloom video creation workspace: text-to-video under the text-to-video tab, and reference-to-video under the image-to-video tab, since it takes image input. The credit cost for your exact duration and resolution is shown before you confirm a run, so nothing bills by surprise.

Verdict

With price and specs identical, this is purely a question of what you are protecting. Protecting a subject across shots: reference-to-video, every time. Protecting your time on a one-off shot with no recurring cast: text-to-video. Anyone building a series, a character-led channel or a product catalogue should default to reference-to-video and treat text-to-video as the tool for filler and exploration.

Frequently asked questions

Do HappyHorse 1.1 Reference-to-Video and Text-to-Video cost different amounts?

No. VdoBloom prices all three HappyHorse 1.1 modes from one shared credit table. A 5-second 720p clip is 45 credits in either mode, and a 15-second 1080p clip is 173 credits in either mode.

How many reference images should I upload?

Two or three that agree with each other works better than one. Consistent lighting and angle across your references gives the model a clearer read on the subject than a single hero shot paired with mismatched extras.

Can both modes output ultrawide 21:9?

Yes. The 1.1 revision added 21:9 and 9:21 alongside 4:5 and 5:4, and all nine ratios are available in both text-to-video and reference-to-video. The 1.0 HappyHorse models only cover the five core ratios.

Is reference-to-video the same as image-to-video?

No. Image-to-video animates your upload as the literal opening frame. Reference-to-video uses your uploads as a guide to the subject’s identity, then generates an entirely new prompted scene around that subject.

What does it cost to try these?

New VdoBloom accounts get 10 free credits with no card, which is enough to explore the interface but not a full 4-second HappyHorse run at 36 credits. Subscriptions start at $15 a month for 300 credits, and one-time credit packs start at $2.49 for 75 credits that never expire.

Ready to try it?

Create your first AI video in minutes β€” no credit card required.

Start Creating Free β†’