Guide7 min readSeptember 13, 2026

How to Use HappyHorse Reference-to-Video for Consistent Characters

A practical guide to HappyHorse Reference-to-Video on VdoBloom: how 1 to 9 reference images keep a subject consistent, real aspect ratios in 1.0 vs 1.1, exact credit costs, and how to build a reference set that holds.

HappyHorse Reference-to-Video takes 1 to 9 reference images of a subject and generates a brand new prompted scene around that subject — 4 to 15 seconds at 720p or 1080p, from 46 credits (about $2.30) to 295 credits (about $14.75) per clip. It is not image animation. Your uploads are never used as the opening frame; they tell the model who the character is, what the product looks like and which style to hold, and everything you see in the output is generated fresh from your prompt.

Reference-to-video versus image-to-video

This is the distinction that decides whether the model works for you. HappyHorse Image-to-Video accepts exactly one still and moves it: the first frame of the clip is your picture, and the model’s job is to add motion without breaking what is already there. Reference-to-video accepts up to nine stills and moves on. It reads them as identity information — face, silhouette, packaging, palette — and then builds a scene your prompt describes, in a location and a pose that never existed in any of your uploads.

The practical payoff is consistency across shots. One reference set can drive a dozen generations, and the same character shows up recognizable in each one instead of drifting into a stranger by clip three. That is what turns a pile of AI clips into something that reads as a series, an episodic ad, or a product campaign rather than a slot machine.

The cost of that power is control over the exact frame. If you have a photograph you love and you want that photograph to move, reference-to-video is the wrong mode — it will produce something adjacent, not your image in motion. Because reference-to-video works from images of a subject, you must have the consent of any real person whose photo you use as a reference.

Real specs, 1.0 and 1.1

Both versions sit in the HappyHorse family at a 4.8 catalog rating and in the MODERATE content tier — standard filtering. They differ in exactly one row.

SettingReference-to-Video 1.0Reference-to-Video 1.1
Reference images1 to 91 to 9
Durations4 to 15 seconds, one-second steps4 to 15 seconds, one-second steps
Resolutions720p, 1080p720p, 1080p
Aspect ratios16:9, 9:16, 1:1, 4:3, 3:416:9, 9:16, 1:1, 4:3, 3:4, 4:5, 5:4, 21:9, 9:21
Content tierMODERATEMODERATE
Catalog rating4.84.8
Credit costSee table belowIdentical to 1.0

So the version choice is simple: if your deliverables ship in the five standard ratios, 1.0 covers them and costs the same. If you need feed-native 4:5, its 5:4 mirror, ultrawide 21:9, or the 9:21 extreme vertical, that is 1.1 territory. There is no price penalty for choosing the newer build, and no capability penalty for staying on 1.0 within the five core ratios.

Note that the aspect ratio picker is genuinely live here, unlike in image-to-video where the output simply follows the uploaded still. Because the scene is generated rather than animated, you choose the framing, and the same reference set can produce a 21:9 establishing shot and a 9:16 vertical cut of the same character.

Exact credit cost per tier

Credits convert at $0.05 each — the Lite plan is $15 for 300 credits. Pricing is by resolution and duration only; the number of reference images you attach does not change the cost.

Duration720p credits720p cost1080p credits1080p cost
4 seconds46$2.3079$3.95
5 seconds58$2.9099$4.95
7 seconds81$4.05138$6.90
8 seconds92$4.60158$7.90
10 seconds115$5.75197$9.85
12 seconds138$6.90236$11.80
15 seconds173$8.65295$14.75

Series work multiplies these numbers, so plan the series rather than the shot. Six 8-second clips at 720p is 552 credits, about $27.60; the same six at 1080p is 948 credits, about $47.40. The sensible pattern is to lock the reference set and the prompt structure at the 46-credit 4-second 720p tier, then spend once at the length and resolution you actually ship. A new account’s 10 free credits will not cover a HappyHorse run, so check the pricing page first — one-time credit packs start at $2.49 for 75 credits and never expire, and paid plans download watermark-free with commercial rights.

How to build a reference set that holds

The output quality here depends more on your uploads than on your prompt wording. A single flattering headshot is the most common mistake; it gives the model one angle and one lighting condition to infer an entire person from.

  • Use three to six images, not one and not nine. Coverage helps, but a large set of near-identical frames adds nothing and a set with contradictory looks pulls the identity apart.
  • Vary angle, not identity. Front, three-quarter and profile of the same subject, same haircut, same era. Photos of a person five years apart are, to the model, two different people.
  • Keep lighting neutral where you can. Heavy colored light bakes into the learned identity and then fights whatever lighting your prompt asks for.
  • Crop tight on what must stay consistent. For a product, include one clean shot of the packaging text and one of the silhouette. For a character, faces should be large in frame.
  • Drop anything blurry or low-resolution. One soft reference is worse than one fewer reference.

On the prompt side, describe the scene and the action, not the subject. The references already carry appearance; restating hair color and outfit in prose only invites the model to reinterpret them. Write the setting, the camera move, the lighting and one clear action — for example: she walks through a sunlit market street, handheld camera follows from behind at shoulder height, warm late-afternoon light, she glances back over her shoulder once.

A worked example: one character, three placements

  1. Assemble four references of the character — front, three-quarter left, three-quarter right, and a mid-body shot — generated in AI image generation or shot yourself.
  2. Open the image-to-video workspace and select HappyHorse Reference-to-Video 1.1.
  3. Test the set: 4 seconds, 720p, 16:9, plain prompt. 46 credits, $2.30. Check that the face survives the generation before spending anything more.
  4. Hero cut: 8 seconds, 1080p, 21:9. 158 credits, $7.90.
  5. Feed cut: 8 seconds, 1080p, 9:16, same references, prompt rewritten for a tighter framing. 158 credits, $7.90.
  6. Square cut: 6 seconds, 1080p, 1:1. 119 credits, $5.95.

Total: 481 credits, about $24.05, for three placement-native cuts of one consistent character. Doing the same job with image-to-video would mean producing three separate source stills that already match each other — which is the harder half of the problem, and exactly the half reference-to-video removes.

When to pick a different model

Choose HappyHorse Image-to-Video when you have the exact frame you want and only need motion added to it — same 4 to 15 second range, same credit table, one image in, your image on screen.

Choose HappyHorse Text-to-Video when there is no subject to keep consistent. Prompt-only generation avoids the effort of building a reference set, and it shares the same durations, resolutions and prices; 1.1 carries the same nine aspect ratios.

Choose a dedicated effect tab under video creation when the motion is a known, repeatable one rather than a bespoke scene. A tuned template beats prose instructions for standard moves, and generally costs less per clip.

Frequently asked questions

How many reference images can I attach?

Between 1 and 9. The endpoint validates the count and rejects a request with zero or more than nine. Three to six varied, sharp images of one subject is the practical sweet spot.

Do extra reference images cost more credits?

No. Cost depends only on resolution and duration — from 46 credits for 4 seconds at 720p to 295 credits for 15 seconds at 1080p. A nine-image set costs the same as a one-image set at the same length.

Will my reference image appear as the first frame?

No, and that is the defining behavior of this mode. The clip is generated from your prompt with the subject’s identity carried over. If you need your exact picture on screen, use image-to-video instead.

Which version should I pick?

Version 1.0 if you ship in 16:9, 9:16, 1:1, 4:3 or 3:4. Version 1.1 if you need 4:5, 5:4, 21:9 or 9:21. Durations, resolutions and credit costs are identical, so there is no budget reason to prefer one.

Can I use photos of a real person as references?

Only with that person’s consent. The MODERATE content tier blocks prohibited categories at generation time, but likeness permission is your responsibility. Explicit and illegal content stays blocked regardless of references.

How consistent is the character really?

Consistent enough for episodic short-form when the reference set is strong and the prompts stay in a similar visual register. Expect small variation between clips, which is why teams usually generate a batch and select, rather than assuming the first render of each shot is final.

Ready to try it?

Create your first AI video in minutes — no credit card required.

Start Creating Free →