Wan AI 2.7 Text-to-Video: Text to Video Online
Bottom line
Wan 2.7 generates picture and sound together -- up to 15 seconds at 1080p, from roughly 6.4 credits/second. Type a prompt into the Text to Video tool and Wan AI 2.7 Text-to-Video renders it on VdoBloom credits — the cost is shown before you commit, and no Wan AI API key is needed.

Wan 2.7 Text-to-Video addresses a common gap in text-to-video models directly: audio support. It produces sound alongside the picture, so a render can move from prompt toward publishable without a separate audio pass. Clips run 5, 10 or 15 seconds at 720p or 1080p, in five aspect ratios -- 16:9, 9:16, 1:1, 4:3 and 3:4 -- covering widescreen, vertical, square and both portrait formats.
As the newest text-to-video generation in Alibaba's Wan line on VdoBloom, 2.7 is also the cheaper option: roughly 6.4 credits/second at 720p and 9.5-9.6 credits/second at 1080p, against Wan 2.6's roughly 7-8 and 12-12.7 credits/second for the same resolutions. Alibaba has not released Wan 2.7's weights publicly, so a hosted platform is the practical way to run it.
Content filtering follows the family pattern: a HOT label, VdoBloom's permissive tier, meaning prompts that stricter STRICT-rated models decline generally render here. Every render is priced in credits with the amount shown up front, so testing whether the generated audio fits your use case costs exactly what the composer says it will, and a failed job is refunded automatically.
What it costs on VdoBloom
Wan 2.7 Text-to-Video bills per second: roughly 6.4 credits/second at 720p, 9.5-9.6 credits/second at 1080p. Worked examples: 5s/720p = 32 credits, 10s/720p = 64 credits, 15s/720p = 95 credits; 5s/1080p = 48 credits, 10s/1080p = 95 credits, 15s/1080p = 143 credits.
Wan 2.7 vs Wan 2.6 vs Wan 3.0
All three generate video from text on VdoBloom, but they scale differently. Wan 2.6 Text-to-Video: 5/10/15s, 720p/1080p, no audio, roughly 7-8 to 12-12.7 credits/second. Wan 2.7 Text-to-Video: same duration range, adds audio and two more aspect ratios, and is cheaper per second (roughly 6.4 to 9.5-9.6 credits/second). Wan 3.0: every second from 2 to 30, an added 480p tier, native audio, and the same t2v/i2v model handles both input types -- at a matching 5-second 720p clip it costs the same 32 credits as Wan 2.7, but the ceiling is double the duration. For a clip under 15 seconds with no need to switch to an image start, 2.7 is the simpler pick; for anything longer or where you want one model for both text and image starts, Wan 3.0 is built for that.
What creators use Wan AI 2.7 Text-to-Video for
- Generate a 15-second 1080p clip with sound in one pass for fast social publishing.
- Produce square 1:1 and vertical 9:16 versions of one concept for different placements without cropping.
- Create ambient scenes where the generated audio supplies atmosphere without a separate editing pass.
- Draft a concept at 720p, confirm the audio and motion work, then re-run at 1080p for delivery.
How to use Wan AI 2.7 Text-to-Video on VdoBloom
- 1Open the Text to Video tool and select Wan AI 2.7 Text-to-Video from the model picker.
- 2Write your prompt.
- 3Choose your duration (5–15s), quality (720p/1080p), aspect ratio, review the credit cost, and generate.
Frequently asked questions
How much does Wan 2.7 Text-to-Video cost on VdoBloom?
Roughly 6.4 credits/second at 720p, 9.5-9.6 credits/second at 1080p. A 5-second 720p clip is 32 credits, 10 seconds is 64, 15 seconds is 95; at 1080p those are 48, 95 and 143 credits.
Does Wan 2.7 Text-to-Video generate sound as well as picture?
Yes -- audio support is the defining feature of this generation. Renders come with generated sound rather than silent video, removing a post-production step for social-first content.
Is Wan 2.7 open source?
No -- Alibaba has not released Wan 2.7's weights, unlike the earlier Wan 2.1/2.2 generations. It's accessed through hosted platforms like VdoBloom, not downloaded and run locally.
Can I run Wan 2.7 locally instead of using an API?
No -- since the weights haven't been published, there's no local ComfyUI or GitHub path for Wan 2.7 the way there is for Wan 2.2. VdoBloom's hosted access is how you get to run it without Alibaba's own paid API.
What durations and resolutions does Wan 2.7 support?
Clips of 5, 10 or 15 seconds at 720p or 1080p. The 15-second option paired with 1080p makes it suitable for a complete scene beat, not just a short loop.
Which aspect ratios can I generate in?
Five: 16:9, 9:16, 1:1, 4:3 and 3:4. That covers widescreen, vertical, square and both portrait formats, so one prompt can be re-run natively for each placement instead of cropping.
How strict is Wan 2.7's content filtering?
It carries the HOT label, VdoBloom's most permissive tier. Filtering is notably lighter than on STRICT-rated models like Runway or Veo, though baseline platform-wide rules still apply to every generation.
How does Wan 2.7 compare to Wan 2.6 Text-to-Video?
Wan 2.7 is cheaper (roughly 6.4 credits/second at 720p vs 7-8 for 2.6), generates audio, and adds two aspect ratios (1:1 and 4:3). Both cap at 15 seconds and share the same 720p/1080p tiers, so 2.7 is generally the better default unless a workflow is already built on 2.6.
How does Wan 2.7 Text-to-Video compare to Wan 3.0?
At a matching 5-second 720p clip, the price is identical -- 32 credits on both. Wan 3.0 extends much further, though: durations up to 30 seconds, an added 480p tier, and unified text-and-image input in one model, versus 2.7's 15-second cap and text-only mode. If you need longer clips or want to switch between text and image starts without changing models, Wan 3.0 is the option; for a straightforward 15-second-or-under text prompt, 2.7 covers it at the same starting price.
What are good prompt-writing habits for Wan 2.7 Text-to-Video?
Since the model also generates audio, describe sound-relevant elements (dialogue, ambient noise, music mood) alongside the visual action -- not just what's on screen. For longer 15-second clips, write in terms of a shot with a beginning and an end rather than a single static description, since the extra length is meant for an actual scene beat.