HappyHorse 1.1 Image-to-Video vs Reference-to-Video: Same Price, Different Job
Both HappyHorse 1.1 modes cost identical credits, so the choice is about what your upload means. Real pricing, the nine-ratio difference, and which mode wins each use case.
HappyHorse 1.1 Image-to-Video and HappyHorse 1.1 Reference-to-Video cost exactly the same on VdoBloom β both draw from one shared price matrix, 36 credits for a 720p 4-second clip up to 173 credits for 1080p at 15 seconds β so the choice between them is purely about what your upload means: Image-to-Video treats your picture as the literal first frame, while Reference-to-Video treats it as a description of a subject and builds a brand-new scene around it. Get that distinction right and the two models stop competing; they do different jobs at the same price.
Both sit in the HappyHorse 1.1 line, both rate 4.8 in the catalog, and both carry the MODERATE content label β the standard middle tier, not the strict one. Durations run 4 to 15 seconds at every whole second on both, and both output 720p or 1080p. That is where the identical part ends.
The one specification that actually differs
Aspect ratio control. Image-to-Video 1.1 has none: the output frame follows the still you uploaded, because that still is the opening frame and cropping it would defeat the point. Reference-to-Video 1.1 exposes the full nine-ratio set β 16:9, 9:16, 1:1, 4:3, 3:4, 4:5, 5:4, 21:9 and 9:21 β because it is composing a new shot rather than extending an existing one.
That is a bigger practical difference than it sounds. With Reference-to-Video you can establish a character once and then generate a 21:9 ultrawide establishing shot and a 9:16 vertical cut of that same character from the identical reference set. With Image-to-Video, a vertical deliverable requires a vertical source image.
| Spec | HappyHorse 1.1 Image-to-Video | HappyHorse 1.1 Reference-to-Video |
|---|---|---|
| What your upload becomes | The literal opening frame | A guide to subject identity and look |
| Durations | 4β15s, every whole second | 4β15s, every whole second |
| Resolutions | 720p, 1080p | 720p, 1080p |
| Aspect ratios | Follows source image | 16:9, 9:16, 1:1, 4:3, 3:4, 4:5, 5:4, 21:9, 9:21 |
| Content tier | MODERATE | MODERATE |
| Catalog rating | 4.8 | 4.8 |
| Offered inside photo-effect tabs | Yes | No β excluded by design |
| 720p, 5s | 45 credits ($2.25) | 45 credits ($2.25) |
| 720p, 10s | 90 credits ($4.50) | 90 credits ($4.50) |
| 720p, 15s | 135 credits ($6.75) | 135 credits ($6.75) |
| 1080p, 5s | 58 credits ($2.90) | 58 credits ($2.90) |
| 1080p, 10s | 116 credits ($5.80) | 116 credits ($5.80) |
| 1080p, 15s | 173 credits ($8.65) | 173 credits ($8.65) |
Dollar figures use the Lite-plan rate of $15 for 300 credits, so 1 credit = $0.05. The full ladder starts at 36 credits for 720p/4s and 47 credits for 1080p/4s, and climbs in clean steps: 720p bills at 9 credits per second, 1080p at roughly 11.5 credits per second.
Why the identical pricing is the important fact
On most platforms a reference-conditioned model carries a premium over plain animation, which pushes people to pick the cheaper one for budget reasons rather than fit reasons. Here the two share one price matrix at the backend level β the same pricing keys resolve for both modes, and for Text-to-Video 1.1 as well β so the cost of choosing wrong is not money, it is a wasted generation.
Which means you should choose on intent alone. Do you want this exact picture to move? Image-to-Video. Do you want this subject to appear in a shot you are about to describe? Reference-to-Video.
A behaviour worth knowing before you go looking for it
Reference-to-Video 1.1 is deliberately excluded from VdoBloomβs photo-effect tabs β the template-driven flows where you upload a portrait and pick a preset motion. The reason is structural: those tabs are built around animating the uploaded photo, and their presets promise to preserve facial features, clothing and lighting exactly. A model that re-generates the subject inside a new scene is the wrong tool for that promise, so it is kept out of the roster. Image-to-Video 1.1 does appear there.
So if you are hunting for Reference-to-Video inside an effect tab and cannot find it, that is intended, not a bug. You use it from the main image-to-video generation flow, where you supply the references and write the scene yourself.
Verdict by use case
Animating a photograph, render or AI-generated still you already like: Image-to-Video. Any picture you made elsewhere β including anything from image generation β becomes a video this way, and the 15-second ceiling gives room for slow drifts and reveals that 5-second models simply cannot fit.
A recurring character or mascot across multiple clips: Reference-to-Video. This is the consistency problem it was built for. Establish the subject once, then generate shot after shot without the character drifting into a stranger by the third clip.
One product in many settings: Reference-to-Video. Studio, lifestyle, outdoor β all generated from a single reference set, all recognisably the same object. Image-to-Video can only move the photo you already have.
Matched ultrawide and vertical cuts of one subject: Reference-to-Video, and it is not close. Nine aspect ratios against zero framing control makes this a one-model decision.
Product photography that must stay pixel-accurate: Image-to-Video. When a client approved that frame, re-generating the product from references is a risk you do not need to take. Animate the approved image instead.
A worked example
You have a brand mascot and need a three-clip sequence: a 21:9 establishing shot, a 9:16 vertical for stories, and a square 1:1 for the feed. On Reference-to-Video 1.1 you upload the mascot references once and run three generations at 720p/8s β 72 credits each, 216 credits total, $10.80 β each with its own aspect ratio and prompt. The mascot stays recognisable across all three.
Trying the same brief on Image-to-Video would mean producing three separate source images at three different aspect ratios first, then animating each one β more steps, more drift risk, and the same 216 credits for the video half of the job. That is the whole argument for the reference mode in one paragraph.
Conversely, if you already have one approved 9:16 portrait and simply want it to breathe for eight seconds, Image-to-Video at 720p/8s costs the same 72 credits and gets you there in a single step. If you need generation with no upload at all, the third member of the line, HappyHorse 1.1 Text-to-Video, shares the same duration range, the same resolutions and the same nine ratios.
Whichever mode you use, if the subject is a real person you must have that personβs consent before uploading their photo as a first frame or as a reference.
Frequently asked questions
Do Image-to-Video and Reference-to-Video 1.1 cost different amounts?
No. Both resolve to the same HappyHorse 1.1 price matrix: 36 to 135 credits at 720p and 47 to 173 credits at 1080p, depending on duration. Text-to-Video 1.1 shares that matrix too.
Can Reference-to-Video use more than one reference image?
Yes β multi-image reference semantics are exactly what separates it from single-frame animation, and they are the reason it is kept out of the one-photo effect tabs.
Why can I not choose an aspect ratio on Image-to-Video 1.1?
Because your uploaded still is the opening frame, the output inherits its framing. Aspect-ratio selection only appears on the reference and text modes, which compose a new shot from scratch.
How does HappyHorse 1.1 differ from the original 1.0 line?
The 1.0 reference and text models cover five aspect ratios (16:9, 9:16, 1:1, 4:3, 3:4); 1.1 adds 4:5, 5:4, 21:9 and 9:21. Duration and resolution ranges are the same on both generations, and 1.0 remains in the catalog alongside 1.1.
What does a 15-second 1080p clip cost in dollars?
173 credits, which is $8.65 at the Lite rate of $15 for 300 credits. Annual billing brings the effective per-credit price down β see pricing for the current plans and the one-time packs that never expire.
Is either model heavily filtered?
Both are MODERATE, the standard middle tier rather than the strict one. Compare the full spec sheets at Image-to-Video 1.1 and Reference-to-Video 1.1.
Ready to try it?
Create your first AI video in minutes β no credit card required.
Start Creating Free β