HappyHorse Image-to-Video vs Reference-to-Video: Which One Fits Your Shot
Both HappyHorse 1.0 modes cost identical credits on VdoBloom. The real split is inputs: one image and no reframing versus up to nine references and full aspect ratio control.
HappyHorse Image-to-Video animates one photo you already have, while HappyHorse Reference-to-Video uses up to nine images as a description of a subject and then generates a brand-new scene around it β both cost identical credits on VdoBloom, so the choice is purely about whether you want your picture moved or your subject re-cast. They sit in the same HappyHorse 1.0 family, share the same 4β15 second duration ladder, the same 720p and 1080p resolutions, and the same credit table. Everything that actually separates them lives in the inputs.
The one-line difference
Image-to-Video takes exactly one image. The backend rejects anything else β the service validates that the upload list contains a single URL β and that image becomes the literal opening frame of the clip. Whatever is in it stays in it: the framing, the background, the lighting, the crop. Your prompt only describes the motion.
Reference-to-Video takes between one and nine images, and the prompt is mandatory rather than optional. Those images are not the first frame. They are evidence about who or what the subject is, and the model builds a new prompted scene that keeps the subject recognisable. Change the prompt and you change the location, the wardrobe, the camera and the action, while the face or the product stays consistent.
Specs side by side
| Spec | HappyHorse Image-to-Video 1.0 | HappyHorse Reference-to-Video 1.0 |
|---|---|---|
| Image inputs | Exactly 1 | 1 to 9 reference images |
| Prompt | Optional (up to 5,000 characters) | Required (up to 5,000 characters) |
| Durations | 4β15s, every whole second | 4β15s, every whole second |
| Resolutions | 720p, 1080p | 720p, 1080p |
| Aspect ratio control | None β inherited from your image | 16:9, 9:16, 1:1, 4:3, 3:4 |
| Content tier | MODERATE | MODERATE |
| Family rating on VdoBloom | 4.8 | 4.8 |
| Available in effect tabs | Yes | No β excluded by design |
That aspect-ratio row is the spec people trip over. Image-to-Video exposes no aspect ratio setting at all, because the frame you uploaded already decides it. If you need a 9:16 vertical from a 16:9 photo, you crop the photo first β the model will not reframe it for you. Reference-to-Video, by contrast, treats framing as a free parameter, so one set of references can produce a widescreen establishing shot and a vertical cut without touching the source files.
Credit cost: identical, so it is not a tiebreaker
Both modes bill from the same per-second matrix. At the Lite rate of 300 credits for $15, one credit is $0.05.
| Resolution & duration | Credits | Dollar cost | Typical use |
|---|---|---|---|
| 720p, 4s | 46 | $2.30 | Cheapest test run |
| 720p, 5s | 58 | $2.90 | Standard social beat |
| 720p, 8s | 92 | $4.60 | Two-action sequence |
| 720p, 10s | 115 | $5.75 | Full camera move |
| 720p, 15s | 173 | $8.65 | Longest clip available |
| 1080p, 4s | 79 | $3.95 | Delivery-grade short cut |
| 1080p, 5s | 99 | $4.95 | Most common finished clip |
| 1080p, 8s | 158 | $7.90 | Ad-length hero shot |
| 1080p, 10s | 197 | $9.85 | Product reveal |
| 1080p, 15s | 295 | $14.75 | Maximum-length finished piece |
Two practical consequences. First, iterate at 720p and only finish at 1080p β the 1080p premium is roughly 70 percent at every length, and at 15 seconds that is 122 credits of difference per attempt. Second, because both modes cost the same, there is no financial reason to force a reference workflow through image animation or the reverse. Pick the one that matches the job. Full plan and pack pricing is on the pricing page, and one-time credit packs never expire, which suits the bursty way most people use 15-second renders.
When Image-to-Video is the right call
Choose HappyHorse Image-to-Video when the picture is already the shot. A product photograph you art-directed, a poster frame from a shoot, a still you generated and then retouched β anything where the composition is the asset and you only want it to breathe. It is also the mode that plugs into VdoBloom effect tabs, so it sits behind the photo-motion workflows on the Image to Video tab. Reference-to-Video is deliberately kept out of those effect rosters, because multi-image reference semantics are the wrong fit for an interface whose whole promise is animate this exact photo.
Prompting is different too. Because the prompt is optional here, a short motion-only instruction usually beats a scene description. Write what moves, not what exists: slow push-in, fabric drifting in the wind, steam rising from the cup, subject turns her head toward camera. Describing the room again invites the model to redecorate a frame you already liked.
When Reference-to-Video wins
Choose HappyHorse Reference-to-Video whenever the same subject needs to appear in more than one shot. A recurring character in an episodic series, a mascot, a single product photographed once and then placed into studio, street and kitchen settings. Feeding three to five references from different angles gives the model far more to hold onto than one frontal image, which is exactly what the nine-image ceiling is for.
Here the prompt carries the whole scene, because nothing else does. Name the subject, then the setting, action, camera and light: the same woman in the reference images, walking through a neon-lit arcade at night, handheld camera tracking behind her, shallow depth of field. Then set the aspect ratio explicitly rather than accepting the default.
One rule applies to both modes and is not negotiable: if the photo or reference set shows a real person, you must have that personβs consent before generating video of them.
A worked example
Say you have shot one bottle of cold-brew coffee and need a week of content. The mistake is to run Image-to-Video ten times on the same still and get ten near-identical clips with slightly different condensation.
Better: run Image-to-Video once at 1080p for 5 seconds (99 credits, $4.95) to get the hero shot moving from your art-directed frame. Then take four photos of the bottle from different angles as a reference set and run Reference-to-Video three times at 720p for 5 seconds (58 credits each, $2.90) to place the same bottle on a cafe counter, on a picnic blanket and in a gym bag, generating the 9:16 cut directly by setting the aspect ratio. Total: 273 credits, about $13.65, for four genuinely distinct clips instead of one shot repeated.
Verdict by use case
One great photo, one clip: Image-to-Video, every time. One subject, many scenes: Reference-to-Video, no contest. You need a specific aspect ratio the source image does not have: Reference-to-Video, because Image-to-Video has no reframing control. You want an effect-tab workflow: Image-to-Video, since the reference variants are excluded from those rosters. No image at all: neither β use HappyHorse Text-to-Video, which shares the identical price table, or start on the free generator with the 10 credits every new account gets.
Frequently asked questions
Do Image-to-Video and Reference-to-Video cost the same on VdoBloom?
Yes. Both draw from the same HappyHorse 1.0 per-second matrix β 46 credits for 720p at 4 seconds up to 295 credits for 1080p at 15 seconds. Mode does not change price; resolution and duration do.
How many images can I upload to each?
Image-to-Video accepts exactly one image and rejects any other count. Reference-to-Video accepts one to nine reference images, and more angles generally means better subject consistency.
Why can I not pick an aspect ratio on Image-to-Video?
Because your uploaded image is the first frame, its own dimensions define the output. Crop the source before uploading if you need a different shape, or switch to Reference-to-Video, which offers 16:9, 9:16, 1:1, 4:3 and 3:4.
Is the prompt required?
Only for Reference-to-Video, where it defines the entire generated scene. On Image-to-Video the prompt is optional, and when you do write one it should describe motion rather than restate the contents of the frame.
What about the 1.1 versions?
HappyHorse 1.1 covers the same 4β15 second and 720p/1080p range, adds the 21:9, 9:21, 4:5 and 5:4 ratios on its text and reference modes, and is priced from a separate, cheaper matrix. The 1.0 models remain in the catalog for teams with prompt libraries already tuned against them.
Ready to try it?
Create your first AI video in minutes β no credit card required.
Start Creating Free β