How to Use HappyHorse 1.1 Text-to-Video: Nine Aspect Ratios, 4-15 Seconds
A working guide to HappyHorse 1.1 Text-to-Video on VdoBloom: the nine-ratio framing set including 21:9 ultrawide, per-second durations, the exact credit cost of every tier, how it compares to 1.0, and prompt structure that holds up at 12 seconds.
HappyHorse 1.1 Text-to-Video turns a written prompt into a 4 to 15 second clip at 720p or 1080p, and its stand-out feature is framing: nine aspect ratios including true 21:9 ultrawide and its 9:21 vertical mirror, at 36 credits ($1.80) for a 4-second 720p run rising to 173 credits ($8.65) for 15 seconds at 1080p. If you have been cropping widescreen output to fake a cinematic letterbox, this is the model that removes that step.
What it is genuinely best at
Two things separate HappyHorse 1.1 Text-to-Video from the crowded middle of the text-to-video field, and neither is raw image quality.
The first is the framing spread. Most generators ship 16:9, 9:16 and 1:1 and stop there. This one adds 4:3 and 3:4, the 4:5 and 5:4 feed-optimised pair that Instagram and Facebook actually favour, and then 21:9 and 9:21. Ultrawide is rare enough in AI video that people routinely render 16:9 and crop, losing a third of their vertical resolution in the process. Here you compose for the ratio from the start.
The second is duration granularity. Durations run from 4 to 15 seconds at every whole second β not a choice of 5 or 10. That sounds cosmetic until you are cutting to music or matching a voiceover, where a 7-second shot that fits beats a 5-second shot you have to stretch or a 10-second shot you have to trim. It also means a slow camera push or a two-beat action has somewhere to breathe, which 5-second-capped models cannot offer at any price.
What it is not: an audio model. HappyHorse generates picture only, with no native speech or sound design. Plan for a separate audio pass.
Specs and exact credit cost
Cost is a function of duration and resolution only. Aspect ratio is free, which is why composing for 21:9 costs nothing extra. Converting at the Lite plan rate β $15 for 300 credits, so one credit is $0.05:
| Duration | 720p credits | 720p cost | 1080p credits | 1080p cost |
|---|---|---|---|---|
| 4s | 36 | $1.80 | 47 | $2.35 |
| 5s | 45 | $2.25 | 58 | $2.90 |
| 6s | 54 | $2.70 | 70 | $3.50 |
| 7s | 63 | $3.15 | 81 | $4.05 |
| 8s | 72 | $3.60 | 93 | $4.65 |
| 9s | 81 | $4.05 | 104 | $5.20 |
| 10s | 90 | $4.50 | 116 | $5.80 |
| 11s | 99 | $4.95 | 127 | $6.35 |
| 12s | 108 | $5.40 | 139 | $6.95 |
| 13s | 117 | $5.85 | 150 | $7.50 |
| 14s | 126 | $6.30 | 162 | $8.10 |
| 15s | 135 | $6.75 | 173 | $8.65 |
Aspect ratios: 16:9, 9:16, 1:1, 4:3, 3:4, 4:5, 5:4, 21:9, 9:21. Content filtering is the standard MODERATE tier. The same credit table applies to HappyHorse 1.1 image-to-video and reference-to-video, so switching modes mid-project does not change your budget maths.
One number worth internalising: 720p costs 9 credits per second flat. That makes mental budgeting trivial β seconds times nine. 1080p works out to roughly 11.5 credits per second.
1.1 against the original 1.0
Both versions share durations, resolutions and rating. The differences that matter are framing and price.
| Β | HappyHorse 1.1 T2V | HappyHorse 1.0 T2V |
|---|---|---|
| Aspect ratios | Nine, incl. 21:9 / 9:21 / 4:5 / 5:4 | Five: 16:9, 9:16, 1:1, 4:3, 3:4 |
| 5s at 720p | 45 credits ($2.25) | 58 credits ($2.90) |
| 5s at 1080p | 58 credits ($2.90) | 99 credits ($4.95) |
| 15s at 1080p | 173 credits ($8.65) | 295 credits ($14.75) |
At 15 seconds and 1080p, 1.1 is 41% cheaper than the 1.0 model for the same length and resolution. Unless you have a prompt library tuned tightly against 1.0 behaviour, default to 1.1.
How to prompt it well
Text-to-video has no image to lean on, so every element of the frame has to come from your words. A prompt that reads well to a human often produces mush; a prompt built like a shot list does not. Order it: subject, action, setting, camera, lighting, mood.
Specific habits that pay off here:
- Write the camera explicitly. Slow dolly-in, locked-off, handheld drift, crane up. Left unsaid, the model picks for you, and on longer durations it often picks badly.
- Compose for the ratio you chose. A 21:9 prompt should describe horizontal space β a figure small against a landscape, action entering from frame left. A 9:21 prompt should stack information vertically. Feeding a centred close-up prompt into ultrawide wastes most of the frame.
- Give long clips more beats. For 10 seconds and up, write two or three sequential actions. One-beat prompts at 12 seconds stall or loop.
- Name the light. Golden hour backlight, overcast soft, single hard key, neon practicals. Lighting language moves output quality more than adjective stacking does.
- Drop the negatives. Describe what should be in frame, not what should not.
A worked example
Target: a cinematic product teaser plus a vertical cut, both from prompts alone.
Run one, 21:9, 8 seconds, 1080p (93 credits, $4.65). Prompt: a lone matte-black wireless speaker on a wet concrete plinth in an empty warehouse, water beading on its surface, slow lateral dolly left to right revealing a shaft of light from a high window, cool grey key with a warm rim, dust suspended in the beam, camera settling as the light reaches the product.
Run two, 9:16, 6 seconds, 1080p (70 credits, $3.50). Same subject, restructured vertically: overhead-to-eye-level crane down the full height of the plinth, speaker centred low in frame, light shaft running top to bottom, ending on a static hero framing.
Two placement-ready assets for 163 credits, $8.15. Prototype both at 720p and 4 seconds first β 72 credits for the pair β and only commit to 1080p once the camera move reads right.
When to pick a different model
- You already have the image. Animating an existing still is image-to-video work, not text-to-video.
- You need the same character across many clips. Use HappyHorse 1.1 Reference-to-Video, which costs the same per second but conditions on 1 to 9 reference images so the subject stays consistent. If those references show a real person, you need that personβs consent.
- You need dialogue or synchronised audio. Pick an audio-native or lip-sync model; HappyHorse renders silent video.
- You only need three seconds. The floor here is 4 seconds, and several faster models on the platform start lower and cheaper for quick social cuts.
Start in the text-to-video tab, pick HappyHorse Text-to-Video 1.1 from the model list, set duration and resolution, and the exact credit cost appears before you commit. New accounts get 10 free credits without a card; paid plans export watermark-free with commercial rights, and one-time credit packs start at $2.49 and never expire. Full breakdown on the pricing page.
Frequently asked questions
What aspect ratios does HappyHorse 1.1 Text-to-Video support?
Nine: 16:9, 9:16, 1:1, 4:3, 3:4, 4:5, 5:4, 21:9 and 9:21. The 1.0 version supports only the first five.
How much does a 10-second clip cost?
90 credits at 720p ($4.50) or 116 credits at 1080p ($5.80), using the Lite rate of $0.05 per credit.
Can I generate clips longer than 15 seconds?
Not in a single run β 15 seconds is the hard ceiling and 4 seconds the floor. Longer pieces are built by generating several clips and cutting them together, which is also where reference mode earns its keep.
Does it generate sound?
No. HappyHorse output is silent video. Add music, voiceover or sound design in a separate step.
Is 1080p worth the extra credits?
For final deliverables, yes β it is roughly a 28% premium over 720p at the same length. For iteration it is wasted money; test at 720p and 4 to 5 seconds, then render the winner once.
Ready to try it?
Create your first AI video in minutes β no credit card required.
Start Creating Free β