HappyHorse 1.1 Text-to-Video vs HappyHorse Image-to-Video 1.0
A prompt-only model at 9 credits per 720p second against a photo-animation model at 11.5 - real credit tables, the aspect-ratio catch, and which one actually costs less per usable clip.
HappyHorse 1.1 Text-to-Video builds a clip from your prompt alone and costs 45 credits for 5 seconds at 720p; HappyHorse Image-to-Video 1.0 animates a still you upload and costs 58 credits for the same 5 seconds at 720p. The 1.1 text model is roughly 22 percent cheaper second for second, and it gives you nine aspect ratios, while the 1.0 image model takes its framing from whatever picture you feed it. Pick text-to-video when you have an idea; pick image-to-video when you already have the exact frame.
Two different jobs, not two versions of one
These models are easy to line up because they sit in the same HappyHorse family on VdoBloom, both rated 4.8, both on the MODERATE content label. But they answer different questions. Text-to-Video 1.1 starts from nothing and invents a frame. Image-to-Video 1.0 starts from a frame you already trust β a product photo, a headshot, a render, something you made in the image generator β and puts it in motion.
There is also a generation gap. The 1.1 models are the newer revision of the line; the 1.0 models remain in the catalogue because their behaviour is a known quantity and people have prompt libraries tuned to them. That gap shows up in both price and framing control.
Spec and price, side by side
Every number below comes from the live VdoBloom credit table. Credits convert at 1 credit = $0.05 on the Lite plan ($15 for 300 credits); annual billing lowers the effective rate, and the Pro tier at $49 a month annual for 2,000 credits is cheaper again. See the pricing page for the full ladder.
| Attribute | HappyHorse 1.1 Text-to-Video | HappyHorse Image-to-Video 1.0 |
|---|---|---|
| Input | Prompt only | One still image + prompt |
| Durations | 4β15s, 1-second steps | 4β15s, 1-second steps |
| Resolutions | 720p, 1080p | 720p, 1080p |
| Aspect ratios | Nine: 16:9, 9:16, 1:1, 4:3, 3:4, 4:5, 5:4, 21:9, 9:21 | None selectable β inherited from the source image |
| Content tier | MODERATE | MODERATE |
| 4s at 720p | 36 credits ($1.80) | 46 credits ($2.30) |
| 5s at 720p | 45 credits ($2.25) | 58 credits ($2.90) |
| 10s at 720p | 90 credits ($4.50) | 115 credits ($5.75) |
| 15s at 720p | 135 credits ($6.75) | 173 credits ($8.65) |
| 5s at 1080p | 58 credits ($2.90) | 99 credits ($4.95) |
| 10s at 1080p | 116 credits ($5.80) | 197 credits ($9.85) |
| 15s at 1080p | 173 credits ($8.65) | 295 credits ($14.75) |
The 720p gap is meaningful and the 1080p gap is large. At 720p the 1.1 text model bills 9 credits per second against 11.5 for the 1.0 image model. At 1080p it is about 11.5 per second against roughly 19.7 β a 15-second 1080p clip costs 173 credits on one and 295 on the other, a difference of $6.10 per render.
The honest caveat: compare like with like
That price gap is a generation gap, not an argument that prompting beats uploading. If what you actually need is to animate a photo, the fair comparison is HappyHorse 1.1 Image-to-Video, which is billed from the same shared 1.1 table as the text model: 45 credits for 5 seconds at 720p, 173 for 15 seconds at 1080p. Choosing text-to-video to save credits when your job is animating a specific picture is a false economy β you will burn the savings regenerating a frame you already had.
So read the table this way. Between HappyHorse 1.1 Text-to-Video and HappyHorse Image-to-Video 1.0, the decision is about input, and the price difference is a bonus argument for the newer generation whenever you are choosing between generations at all.
Framing is the second real difference
Text-to-Video 1.1 exposes nine aspect ratios, including true ultrawide 21:9, its vertical mirror 9:21, and the 4:5 and 5:4 feed crops. You ask for the shape you need and the model composes for it.
Image-to-Video 1.0 has no aspect-ratio control at all, and this is by design rather than an omission: the source image decides the frame. If you hand it a square photo you get a square clip. That means your framing decision moves upstream to the picture β crop it before you upload, not after you render. It is a common cause of a wasted run: people upload a 4:3 photo, expect a 9:16 vertical, and get 4:3 motion back.
A worked example
You need a 9:16 vertical teaser for a coffee brand, 8 seconds.
Route one, text-to-video: write the shot β steam rising off a dark espresso in a matte black cup, slow overhead descent, morning window light, shallow depth β select 9:16, 8 seconds, 720p. That is 72 credits, $3.60. If the composition misses, you re-roll for another 72.
Route two, image-to-video: shoot or generate the cup shot first, crop it to 9:16 yourself, upload it, and prompt only the motion. That is 92 credits at 8 seconds and 720p on the 1.0 model, $4.60 β but you already know exactly what the frame looks like, so the re-roll rate is much lower. Two text re-rolls cost more than one image run.
That is the real economics. Per render, text-to-video is cheaper. Per usable render, image-to-video often wins when the composition is non-negotiable, because you are only gambling on motion instead of gambling on motion and composition together.
Prompting each one
For text-to-video, name the subject, the action, the camera move and the light in that order, and give longer clips more than one beat β a 15-second clip described as a single instruction will tend to loop or drift. Be explicit about the shape you want in the aspect-ratio selector rather than asking for it in the prompt.
For image-to-video, do the opposite: do not re-describe what is already in the picture. The upload carries the subject, the palette and the composition. Write only what changes β the camera drifts left, the steam rises, her hair lifts in the breeze. Over-describing the still is the most common reason an animation fights its own source frame.
If your upload is a photograph of a real person, you must have that personβs consent before generating video of them.
Verdict by use case
- Concept exploration, b-roll, abstract motion, ultrawide plates: HappyHorse 1.1 Text-to-Video. Cheapest per render, and the only one of the two with framing control.
- Animating a product shot, headshot or finished render: an image model β and prefer the 1.1 image model over the 1.0 one, since it does the same job from the cheaper shared price table.
- You have a 1.0-tuned prompt library and it works: stay on HappyHorse Image-to-Video 1.0. Predictable behaviour is worth real money; just know you are paying about 50 percent more per 1080p second for it.
- Tight credit budget, long clips: 1.1 Text-to-Video at 720p, 135 credits for the full 15 seconds.
Both live in the VdoBloom video workspace: the text model under the text-to-video tab and the image model under the image-to-video tab. Exact credit cost for your chosen duration and resolution is shown before you confirm, so you never find out the price after the fact.
Frequently asked questions
Which is cheaper, HappyHorse 1.1 Text-to-Video or HappyHorse Image-to-Video 1.0?
The 1.1 text model. It bills 9 credits per second at 720p and about 11.5 at 1080p, against 11.5 and roughly 19.7 for the 1.0 image model. A 10-second 1080p clip is 116 credits versus 197.
Why does HappyHorse Image-to-Video 1.0 have no aspect ratio setting?
Because the source image defines the frame. Crop your photo to the shape you want before uploading; the clip comes back in the same proportions as the still.
Can I use text-to-video to make the image and then animate it?
You can run that two-step workflow, but generate the still in the image generator rather than pulling a frame out of a video, then feed it to an image-to-video model. That gives you a clean, high-resolution first frame.
How long can each model run?
Both support 4 to 15 seconds, selectable at every whole second, at 720p or 1080p. The duration range is identical across the whole HappyHorse line.
What do I get on the free tier?
New accounts get 10 free credits with no card, enough to look around rather than complete a HappyHorse render. Plans start at $15 a month for 300 credits, credit packs start at $2.49 for 75 non-expiring credits, and paid plans download watermark-free with commercial rights.
Ready to try it?
Create your first AI video in minutes β no credit card required.
Start Creating Free β