Guide7 min readSeptember 13, 2026

How to Use HappyHorse 1.1 Reference-to-Video for Consistent Characters

A practical guide to HappyHorse 1.1 Reference-to-Video on VdoBloom: what reference conditioning really does, how many images to upload, the exact credit cost for every duration and resolution, and a worked three-clip example.

HappyHorse 1.1 Reference-to-Video builds a brand-new prompted scene around 1 to 9 reference images, so the same character, product or style stays recognisable from clip to clip β€” at 4 to 15 seconds, 720p or 1080p, in nine aspect ratios, from 36 credits ($1.80) for a 4-second 720p run. It is the model you reach for when consistency across shots matters more than animating one specific photo.

What reference-to-video actually does (and does not do)

The distinction trips people up constantly, so it is worth being blunt about it. Image-to-video takes your upload and treats it as the literal first frame: the output starts on your picture and moves forward from there. Reference-to-video does not do that. Your uploads are treated as a description β€” this is who the subject is, this is what the product looks like, this is the style to hold β€” and the model then generates an entirely new scene from your text prompt with that subject inside it.

The practical payoff is identity persistence. Generate six clips of the same mascot with an ordinary text-to-video model and you get six different mascots. Feed the same reference set into HappyHorse 1.1 Reference-to-Video six times and the subject stays recognisably itself across all six, even though the settings, camera moves and lighting change completely.

What it will not do is guarantee a frame-perfect match. Reference conditioning is a strong steer, not a clone. Fine detail β€” the exact typography on a label, a precise tattoo, the stitching on a logo β€” can drift. If you need the source image preserved pixel-for-pixel in the opening frame, use HappyHorse 1.1 Image-to-Video instead and accept that the camera has to start where your photo is.

One rule that is not optional: if your reference images show a real, identifiable person, you must have that person’s consent before generating video of them.

The real specs and what every run costs

HappyHorse 1.1 prices identically across text, image and reference modes β€” the credit table is shared β€” and the cost is driven only by duration and resolution. Aspect ratio is free. At the Lite plan rate of $15 for 300 credits, one credit is $0.05.

Duration720p credits720p cost1080p credits1080p cost
4s36$1.8047$2.35
5s45$2.2558$2.90
6s54$2.7070$3.50
7s63$3.1581$4.05
8s72$3.6093$4.65
9s81$4.05104$5.20
10s90$4.50116$5.80
11s99$4.95127$6.35
12s108$5.40139$6.95
13s117$5.85150$7.50
14s126$6.30162$8.10
15s135$6.75173$8.65

The rest of the sheet: durations are every whole second from 4 to 15, resolutions are 720p and 1080p, and the aspect ratio set is the full nine β€” 16:9, 9:16, 1:1, 4:3, 3:4, 4:5, 5:4, 21:9 and 9:21. The reference slot accepts between 1 and 9 images; send fewer than one or more than nine and the request is rejected outright. Content filtering sits at the standard MODERATE tier.

Worth noting for anyone migrating: the older 1.0 reference model costs 58 credits for a 5-second 720p clip against 45 on 1.1, and 99 against 58 at 1080p. The newer version is both wider in framing and materially cheaper. There is very little reason to stay on the 1.0 line unless you have prompts tuned hard against it.

Choosing reference images that actually work

Reference quality decides output quality more than prompt wording does. What consistently helps:

  • Vary the angle, not the subject. Three shots of the same character from front, three-quarter and profile teach the model far more than three near-identical front-on frames.
  • Keep the subject dominant in frame. A person occupying 10% of a busy street photo gives the model almost nothing to lock onto. Crop in.
  • Match the lighting you want, roughly. References shot under hard flash will push the generated scene warmer and harder than you may intend.
  • Do not mix two subjects. If half your references are the mascot and half are the product, the model will try to satisfy both and blend them. Run two separate generations.
  • Three to five is the sweet spot. Nine is allowed, but past about five images the marginal gain flattens and contradictory references start fighting each other.

How to prompt it well

Because the subject is carried by the images, the prompt should spend its words on everything except describing the subject in detail. Write the scene, the action, the camera and the light; let the references handle identity. A prompt that re-describes the character in heavy detail tends to override the references rather than reinforce them.

A reliable structure is: subject reference (one short noun phrase) + action + setting + camera move + lighting + ending beat. Give the action a shape with a beginning and an end, especially on longer durations β€” a 12-second clip with a one-beat prompt will loop or stall in the middle.

A worked example

Say you are producing a three-clip teaser for a coffee brand with a recurring barista character. Upload four references of her: front, three-quarter left, profile, and one mid-action shot pouring. Then generate three runs from the same reference set.

  • Clip 1, 21:9, 8s, 1080p (93 credits, $4.65): the barista walks into an empty pre-dawn cafe, houselights flicking on behind her as she passes, slow dolly-in from wide, cool blue window light warming as the lamps rise.
  • Clip 2, 9:16, 6s, 1080p (70 credits, $3.50): the same barista tamps espresso in close-up, steam curling, handheld micro-motion, warm tungsten key from the left, ending as she looks up to camera.
  • Clip 3, 1:1, 5s, 720p (45 credits, $2.25): she slides a finished cup across the counter toward the lens, shallow depth, locked-off camera, ending on the cup settling.

Total: 208 credits, or $10.40 at the Lite rate β€” three placement-ready cuts of a character who looks like the same person in all three. Rendering the vertical cut at 720p instead would bring it to 189 credits.

When to pick a different model

Reference-to-video is the wrong tool more often than its fans admit. Skip it when:

  • You want your exact photo animated. That is image-to-video, not reference mode.
  • You have no subject to keep consistent. A one-off scene with no recurring character is cheaper and simpler as HappyHorse 1.1 Text-to-Video, which costs exactly the same per second but needs no uploads.
  • You need synchronised dialogue or native audio. HappyHorse does not generate speech; a lip-sync or audio-native model is the right path.
  • You are still exploring the idea. Iterate at 720p and 4 to 5 seconds, then commit the winner to 1080p at full length. The difference between a 4-second 720p test and a 15-second 1080p final is 36 credits against 173.

New accounts get 10 free credits with no card, which is not quite enough for a single 4-second run here β€” so try the cheaper models on the free tier first, then come to reference mode once you know what you are building. Paid plans download watermark-free with commercial rights; see pricing for the credit packs, which never expire.

Frequently asked questions

How many reference images can I upload?

Between 1 and 9. The request is rejected if you send none or more than nine. In practice three to five varied shots of a single subject give the best results.

Does the aspect ratio change the credit cost?

No. Cost depends only on duration and resolution. A 21:9 ultrawide 8-second 1080p clip and a 9:16 vertical 8-second 1080p clip both cost 93 credits.

Will the output match my reference photo exactly?

No, and it is not meant to. Reference conditioning preserves identity and style, not pixels. Small details like text on packaging or fine markings can drift between runs.

Can I use a photo of a real person as a reference?

Only with that person’s consent. Likeness rights apply to generated video exactly as they do to photography, and the MODERATE content filter blocks prohibited categories regardless.

What is the cheapest way to test this model?

A 4-second 720p run at 36 credits ($1.80). Lock the references and the prompt structure there before spending 173 credits on a 15-second 1080p final.

Ready to try it?

Create your first AI video in minutes β€” no credit card required.

Start Creating Free β†’