HappyHorse Reference-to-Video vs Text-to-Video: Same Price, Different Job
Same engine, same credits, same aspect ratios. One pins your subject with up to nine reference images; the other invents it from the prompt. Here is when each wins.
HappyHorse Reference-to-Video and HappyHorse Text-to-Video are the same 1.0 engine at the same credit price — the only real difference is that Reference-to-Video lets you attach one to nine images that lock the subject’s identity, while Text-to-Video invents everything from the prompt alone. Identical durations, identical resolutions, identical five aspect ratios, identical cost per second. So the decision is simple: if the same face, mascot or product has to survive from clip to clip, you need references. If it does not, the prompt is faster and gives the model more freedom.
What each one actually takes as input
Text-to-Video requires a prompt between 1 and 5,000 characters, plus duration, resolution and aspect ratio. Nothing else. Every element of the frame — the person, the room, the lens, the light — comes from what you wrote, and from the model’s interpretation of it. Run the same prompt twice with different seeds and you get two different people wearing two different jackets.
Reference-to-Video requires the same prompt, with the same 5,000-character limit and the same settings, and adds a reference image list of one to nine URLs. The prompt still describes the scene, but the references pin down who or what is in it. That is the entire distinction in the API and, honestly, the entire distinction in practice.
Specs and price, side by side
| Spec | HappyHorse Reference-to-Video 1.0 | HappyHorse Text-to-Video 1.0 |
|---|---|---|
| Required input | Prompt + 1–9 reference images | Prompt only |
| Durations | 4–15s in one-second steps | 4–15s in one-second steps |
| Resolutions | 720p, 1080p | 720p, 1080p |
| Aspect ratios | 16:9, 9:16, 1:1, 4:3, 3:4 | 16:9, 9:16, 1:1, 4:3, 3:4 |
| Prompt limit | 5,000 characters | 5,000 characters |
| Content tier | MODERATE | MODERATE |
| Family rating | 4.8 | 4.8 |
| 720p, 5s | 58 credits / $2.90 | 58 credits / $2.90 |
| 1080p, 5s | 99 credits / $4.95 | 99 credits / $4.95 |
| 1080p, 10s | 197 credits / $9.85 | 197 credits / $9.85 |
| 1080p, 15s | 295 credits / $14.75 | 295 credits / $14.75 |
Dollar figures convert at the Lite rate — 300 credits for $15 a month, so one credit is $0.05. The full ladder runs 46, 58, 69, 81, 92, 104, 115, 127, 138, 150, 161 and 173 credits for 720p at 4 through 15 seconds, and 79, 99, 119, 138, 158, 177, 197, 217, 236, 256, 275 and 295 for the same lengths at 1080p. Both modes read from that one table. See the pricing page for plans and for the one-time packs that never expire.
Where Text-to-Video is the better tool
Reach for HappyHorse Text-to-Video when the subject is generic or when you are still exploring. Establishing shots, b-roll, abstract motion, weather, cityscapes, food on a table, anything where no particular face has to come back next week. Because there is no reference set to honour, the model has more latitude and often produces cleaner motion for wide, subject-free scenes.
It is also the cheapest way to find a look. Spend 46 credits on a 720p four-second render, judge the composition, then re-run the winner at 1080p for the length you actually need. Three exploratory 720p tries plus one 1080p five-second final is 237 credits, about $11.85 — less than a single 15-second 1080p render you guessed at.
Prompt it like a shot list rather than a wish. Subject, action, setting, camera, light, in that order: a lone surfer paddling out at dawn, wide shot from the shoreline, slow dolly right, cool blue light with mist on the water. Set the aspect ratio deliberately — 9:16 for vertical feeds, and note that 21:9 is not available here, that is 1.1 territory. Start on the Text to Video tab.
Where Reference-to-Video wins outright
HappyHorse Reference-to-Video earns its keep the moment continuity matters. A recurring host across a series of shorts. A mascot that must look like itself in every episode. One physical product placed into six different settings without six different photo shoots. Text-to-Video simply cannot do this — each generation re-invents the subject, and by clip three your character is a stranger.
The nine-image ceiling exists to be used. One frontal photo is the weakest possible reference; three to five images covering different angles, distances and lighting give the model enough to generalise from. For a product, include a clean pack shot plus a detail crop plus an in-context photo.
Prompt structure changes slightly. Refer to the subject rather than describing it, and spend your words on everything the references cannot carry: the man in the reference images, seated at a diner counter, pouring coffee, medium shot, warm tungsten light, slight handheld drift. Re-describing his hair colour just gives the model two competing instructions.
Because reference workflows frequently involve photographs of real people, one rule is absolute: you must have that person’s consent before generating video from their likeness.
A worked comparison
Imagine a six-part explainer series with a presenter. Doing it with Text-to-Video means writing an extremely detailed description of the presenter and hoping it holds — it will not, and you will burn credits on rejected takes trying.
Doing it with Reference-to-Video means generating or photographing the presenter once, collecting four reference images, and then running six prompts that change only the setting and the action. At 1080p and 8 seconds each, that is 158 credits per clip, 948 credits total, roughly $47.40 for the set — with the same person in all six. Mixing the two is smarter still: use Text-to-Video for the six b-roll cutaways at 720p and 5 seconds (58 credits each, 348 total, $17.40), since nobody needs a consistent cloud.
Verdict
Recurring character, mascot or product: Reference-to-Video, decisively. Scenery, b-roll, abstract or one-off shots: Text-to-Video, which is faster to set up and gives the model room. Early exploration on any project: Text-to-Video at 720p and 4 seconds, the cheapest cell in the table at 46 credits. You already have the exact frame you want moving: neither — that is HappyHorse Image-to-Video, which animates a single upload as the literal first frame at the same cost. New accounts get 10 free credits with no card, and you can try the workflow on the free AI video generator before committing to a plan.
Frequently asked questions
Is Reference-to-Video more expensive than Text-to-Video?
No. Both read from the identical HappyHorse 1.0 credit matrix, from 46 credits at 720p for 4 seconds to 295 credits at 1080p for 15 seconds. Adding reference images costs nothing extra.
How many reference images should I actually supply?
The model accepts one to nine. One works, but three to five covering different angles and lighting gives noticeably steadier subject consistency than a single frontal shot.
Can Text-to-Video keep a character consistent using a very detailed prompt?
Not reliably. Detailed descriptions narrow the range but each generation still re-invents the person. If the same face must return in a later clip, use reference images.
Do both support vertical video?
Yes — both offer 16:9, 9:16, 1:1, 4:3 and 3:4. The wider set that includes 21:9, 9:21, 4:5 and 5:4 belongs to the HappyHorse 1.1 text and reference models.
What content rules apply?
Both sit in the MODERATE tier, which is the standard filtering level: ordinary creative work, fashion, fitness and lifestyle subjects generate normally, while explicit and illegal content is blocked at generation time.
Ready to try it?
Create your first AI video in minutes — no credit card required.
Start Creating Free →