How to Use MiniMax H3 Text-to-Video: Settings, Prompts and Credit Costs
A practical guide to MiniMax H3 Text-to-Video on VdoBloom: 4-15 second clips with native audio, 768P and 2K tiers, exact per-second credit costs, prompting technique and when to pick a cheaper model.
MiniMax H3 Text-to-Video on VdoBloom turns a written prompt into a 4–15 second clip with native audio at 768P or 2K, and it costs 15 credits per second at 768P or 24 credits per second at 2K — so a 5-second 768P clip is 75 credits ($3.75) and a 15-second 2K clip is 360 credits ($18.00). It is the longest single-shot text-to-video model in the VdoBloom catalog, and the only one in its family that carries a 4.9 user rating. This guide covers what it is genuinely good at, its exact settings, the real credit table, how to prompt it, and when a cheaper model is the smarter call.
What MiniMax H3 Text-to-Video is best at
H3 is a cinematic long-form model. Most text-to-video models on the platform cap out at 5, 8 or 10 seconds and ask you to stitch shots together. H3 accepts any whole-second duration from 4 to 15, which means an entire beat — a character walks in, speaks, reacts, and the camera settles — can live in one generation with no cut.
Three things follow from that:
- Audio comes with the clip. H3 generates sound alongside the picture, so ambience and on-screen speech arrive baked in rather than added in an edit.
- It rewards long prompts. A 15-second shot needs 15 seconds of described action. Short prompts produce drifting, aimless motion at long durations.
- It supports 21:9. The text-to-video variant is one of the few models on VdoBloom that offers a true cinemascope ratio, alongside 16:9, 4:3, 1:1, 3:4 and 9:16.
The model carries a MODERATE content label, which puts it in the middle of the VdoBloom filter tiers: it is more permissive than the strict-tier models on swimwear, dance and fashion subject matter, while explicit and illegal content stays blocked outright.
Real specs
| Setting | What MiniMax H3 Text-to-Video accepts |
|---|---|
| Input | Text prompt only (a separate H3 Image-to-Video variant handles photos) |
| Durations | Every whole second from 4 to 15 |
| Quality tiers | 768P and 2K |
| Aspect ratios | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 |
| Audio | Generated with the clip |
| Content tier | MODERATE |
| User rating | 4.9 |
| Where it lives | The Text to Video tab |
Exact credit cost per tier
H3 prices strictly per second, with no minimum beyond the 4-second floor. At the VdoBloom Lite rate of $15 for 300 credits, one credit is $0.05.
| Duration | 768P credits | 768P cost | 2K credits | 2K cost |
|---|---|---|---|---|
| 4s | 60 | $3.00 | 96 | $4.80 |
| 5s | 75 | $3.75 | 120 | $6.00 |
| 6s | 90 | $4.50 | 144 | $7.20 |
| 8s | 120 | $6.00 | 192 | $9.60 |
| 10s | 150 | $7.50 | 240 | $12.00 |
| 12s | 180 | $9.00 | 288 | $14.40 |
| 15s | 225 | $11.25 | 360 | $18.00 |
The rates in between follow the same arithmetic: 15 credits per second at 768P, 24 credits per second at 2K. Durations not listed above (7s, 9s, 11s, 13s, 14s) are all selectable and priced on that line. The 2K tier is a flat 60% premium over 768P at every length, which makes the choice simple: pick 2K when the clip is the finished deliverable, and 768P while you are still testing whether the prompt works at all.
The exact credit figure for the combination you have selected appears in the generate button before you spend anything, so you never have to do this arithmetic in your head. If the number is more than your balance, the pricing options open instead of the job starting. Full plan and credit-pack detail sits on the pricing page.
How to prompt MiniMax H3 well
Because H3 is a long-duration model, the biggest single improvement to your output is writing a prompt with a timeline in it. Four habits that pay off:
- Describe the shot in beats. Write what happens first, then next, then last. A 12-second prompt that reads as one static description will produce one static shot with nowhere to go after second four.
- Name the camera move explicitly. Slow push-in, locked-off wide, handheld tracking shot alongside, low-angle rise. H3 responds to camera language much more reliably than it responds to vague cinematic adjectives.
- Say what you want to hear. Since audio is generated with the picture, describing the soundscape — rain on a metal roof, distant traffic, a single line of dialogue — changes the result. Leaving it unsaid gives you whatever the model infers.
- Set the light. Golden-hour backlight, hard overhead practicals, overcast diffuse daylight. Lighting descriptions carry more weight per word than style name-dropping.
Start any new idea at 4 or 5 seconds and 768P. A failed 5-second test costs 75 credits; a failed 15-second 2K render costs 360. Lock the composition, camera move and lighting cheaply, then re-run the same prompt at the length and resolution you actually need.
A worked example
Say you want a 10-second 21:9 opener for a coffee brand. A weak prompt is: cinematic shot of coffee being poured.
A prompt built for H3 looks more like this: Locked-off 21:9 wide of a marble counter in a narrow city cafe at 7am. Steam rises from a ceramic cup in the foreground. A barista’s hands enter from the right and pour milk in a slow spiral; the camera pushes in gently over six seconds until the crema fills the lower third. Hard morning sun rakes in from a window on the left, dust visible in the beam. Sound: the low hiss of a steam wand, a spoon on porcelain, faint street noise outside.
At 10 seconds and 2K that job costs 240 credits ($12.00). Run it once at 4 seconds and 768P first — 60 credits — to confirm the framing and the light before committing.
When to pick a different model
H3 is a premium tier, and per second it is one of the more expensive text-to-video options in the catalog. Skip it when:
- You need volume, not length. Twenty 5-second social cuts at 768P H3 is 1,500 credits. Faster models such as Seedance 2 Fast exist for exactly that job — its 720p 5-second tier is 50 credits.
- You are animating a photo. The text-to-video variant takes no image input at all. Use MiniMax H3 Image-to-Video instead; it shares the same 4–15 second range, the same 768P/2K tiers and the same per-second pricing, but has no aspect-ratio control because the source image sets the frame. If you upload a photo of a real person, you need that person’s consent before generating.
- You are building an ad or a UGC spot. H3 Text-to-Video is not offered inside the advertisement and viral-video builders. Generate the shot on the Text to Video tab and bring the clip into the ad workflow, or pick a model the ad builder supports.
- You want a different look. Veo 3 Quality and WAN 3.0 sit in the same premium bracket with very different rendering characters. The full list is on the H3 model page and the models directory.
Frequently asked questions
How long can a MiniMax H3 clip be?
Fifteen seconds, in a single generation, at any whole-second length from four upward. There is no stitching and no extend step involved.
Does MiniMax H3 generate sound?
Yes. Audio is produced with the video rather than added afterwards, which is why describing the soundscape in your prompt changes the output.
What does a 5-second MiniMax H3 video cost?
75 credits at 768P ($3.75) or 120 credits at 2K ($6.00), using the Lite-plan rate of $0.05 per credit. Annual billing lowers the effective rate.
Can I use a photo as the input?
Not with this variant. MiniMax H3 Text-to-Video is prompt-only. The H3 Image-to-Video variant is the one that accepts an uploaded image, at identical per-second pricing.
Is there a watermark?
Free accounts get watermarked downloads. Any paid plan downloads watermark-free with commercial rights, and one-time credit packs start at $2.49 for 75 credits that never expire.
Ready to try it?
Create your first AI video in minutes — no credit card required.
Start Creating Free →