Wan AI 2.6 Text-to-Video: Text to Video Online

Text to Video★ 4.3 / 5Flexible content filter

Bottom line

Text-only clips up to 15s at 720p/1080p from Alibaba's Wan 2.6, with VdoBloom's permissive content badge. Type a prompt into the Text to Video tool and Wan AI 2.6 Text-to-Video renders it on VdoBloom credits — the cost is shown before you commit, and no Wan AI API key is needed.

Example output generated with Wan AI 2.6 Text-to-Video

Wan 2.6 Text-to-Video converts a written prompt into finished footage with no image input required. The duration options set it apart from much of the catalog: 5, 10, or a full 15 seconds per clip, at 720p or 1080p. Fifteen seconds is enough for an actual scene beat -- an action with a beginning and an end -- rather than the isolated moment most 5-second models produce, which changes how you should write prompts: think in shots, not stills.

Pricing runs roughly 7-8 credits/second at 720p and 12-12.7 credits/second at 1080p -- the same rates as Wan 2.6 Image-to-Video. That's a bit more than the newer Wan 2.7 Text-to-Video (roughly 6.3-6.4 credits/second at 720p, 9.5-9.6 at 1080p), which also generates audio and adds two extra aspect ratios.

Wan 2.6 carries a HOT content label, VdoBloom's permissive tier, so creative directions that trip filters on stricter models generally pass here. Alibaba released this generation as a commercial model rather than open weights, so VdoBloom's hosted API is the way to run it -- no download or GPU involved. Generation costs are paid in credits and displayed before you run the job, making it straightforward to weigh a 15-second 1080p render against a shorter draft pass.

What it costs on VdoBloom

Wan 2.6 Text-to-Video is priced per second at roughly 7-8 credits/second at 720p and 12-12.7 credits/second at 1080p, edging up slightly at longer durations. Worked examples: 5s/720p = 35 credits, 10s/720p = 75 credits, 15s/720p = 120 credits; 5s/1080p = 60 credits, 10s/1080p = 125 credits, 15s/1080p = 190 credits.

Wan 2.6 vs Wan 2.7 Text-to-Video

Wan 2.7 Text-to-Video is the newer generation on the same dashboard: it costs slightly less per second (roughly 6.4 credits/second at 720p vs 7-8 for Wan 2.6), generates audio alongside the picture, and supports five aspect ratios instead of three. Both cap at 15 seconds and offer the same 720p/1080p tiers. Wan 2.6 remains available for workflows already built around it, but Wan 2.7 is the stronger default for new text-to-video work.

What creators use Wan AI 2.6 Text-to-Video for

  • Write a full 15-second scene beat -- setup, action, resolution -- as a single 1080p clip.
  • Generate stylized b-roll from text descriptions when stock footage falls short.
  • Draft concepts at 720p, then re-run the strongest prompt at 1080p for the final delivery.
  • Prototype creative directions that stricter, STRICT-rated models on VdoBloom decline to render.

How to use Wan AI 2.6 Text-to-Video on VdoBloom

  1. 1Open the Text to Video tool and select Wan AI 2.6 Text-to-Video from the model picker.
  2. 2Write your prompt.
  3. 3Choose your duration (5–15s), quality (720p/1080p), review the credit cost, and generate.

Frequently asked questions

How much does Wan 2.6 Text-to-Video cost?

Roughly 7-8 credits/second at 720p, 12-12.7 credits/second at 1080p (the rate rises slightly with duration). A 5-second 720p clip is 35 credits, 10 seconds is 75, 15 seconds is 120; at 1080p those become 60, 125 and 190 credits.

How long can Wan 2.6 Text-to-Video clips be?

5, 10 or 15 seconds. The 15-second ceiling is among the longer options in VdoBloom's text-to-video lineup and gives a prompt room for a complete action rather than a single moment.

Is Wan 2.6 open source?

No -- unlike the earlier Wan 2.1/2.2 generations that Alibaba released openly, Wan 2.6 launched as a commercial model. It's accessed through hosted platforms like VdoBloom rather than downloaded and run locally.

What does the HOT content label mean for Wan 2.6?

HOT runs across the entire Wan line on VdoBloom -- it's the platform's most permissive filtering tier, so stylized action, combat choreography and darker scene concepts that suit a full 15-second beat will generally render where stricter models refuse. Platform-wide rules on illegal content still apply.

Does Wan 2.6 Text-to-Video generate audio?

No -- it's silent video from text. If you want generated audio in the Wan family, Wan 2.6 Flash (image-to-video) and Wan 3.0 both produce sound natively; Wan 2.7 Text-to-Video also has audio support.

How does Wan 2.6 compare to Wan 2.7 Text-to-Video?

Wan 2.7 is newer, cheaper per second (roughly 6.4 credits/second at 720p vs Wan 2.6's roughly 7-8), generates audio, and offers two more aspect ratios (1:1 and 4:3 alongside 16:9, 9:16 and 3:4). Both share the same 5/10/15-second duration options and 720p/1080p tiers. Wan 2.7 Text-to-Video is generally the better default unless you have a specific reason to stay on 2.6.

Can I write a prompt in a non-English language?

VdoBloom's model specs don't document a language restriction for Wan 2.6, and the underlying model is multilingual by design. If a prompt in your language doesn't render as expected, try an English rewrite to compare results before assuming it's unsupported.

What resolutions does Wan 2.6 Text-to-Video output?

720p and 1080p. A common pattern is drafting concepts at 720p and re-running the strongest prompt at 1080p for the final clip, with the exact credit cost shown before each run.

Who develops the Wan model family?

Wan is Alibaba's video model line. Earlier generations (2.1, 2.2) were released as open weights; 2.5 onward, including 2.6, have shipped as commercial, API-only releases. Version 2.6 is hosted on VdoBloom, where you run it directly with credits.

Related models