Wan AI 2.7 Reference-to-Video: Image to Video Online

Image to Video★ 4.3 / 5Flexible content filter

Bottom line

Feed it reference images or video and Wan 2.7 keeps your subject consistent across 5-10s generations. Upload a starting frame, choose Wan AI 2.7 Reference-to-Video in the Image to Video tool, and animate it on VdoBloom — the credit price appears before every generation.

Example output generated with Wan AI 2.7 Reference-to-Video

Most image-to-video tools treat your upload as frame one and improvise everything after it. Wan 2.7 Reference-to-Video works on a different contract: the images or video you supply act as a persistent reference for the subject -- a face, a product, an art style -- while the prompt controls what actually happens in the shot. That separation is what makes it useful for series work, where the same character has to survive across many separate generations rather than one continuous clip.

Clips come out at 5 or 10 seconds in 720p or 1080p, a shorter range than the sibling Wan 2.7 Image-to-Video model's 5-15 seconds -- R2V trades maximum length for its reference-driven consistency. Because references can be video as well as stills, you can also hand it existing footage as a guide rather than describing a look in text.

Pricing runs roughly 6.4 credits/second at 720p and 9.5-9.6 credits/second at 1080p, the same rate as Wan 2.7 Image-to-Video. It carries VdoBloom's HOT content label, the platform's most permissive filtering tier, and Alibaba has not released Wan 2.7's weights publicly, so VdoBloom's hosted access is the way to run it. The exact credit cost for your chosen duration and resolution is shown before you commit a run -- worth knowing when you're batching a whole set of reference-locked shots.

What it costs on VdoBloom

Wan 2.7 Reference-to-Video bills per second: roughly 6.4 credits/second at 720p, 9.5-9.6 credits/second at 1080p -- the same rate as Wan 2.7 Image-to-Video. Worked examples: 5s/720p = 32 credits, 10s/720p = 64 credits; 5s/1080p = 48 credits, 10s/1080p = 95 credits.

R2V vs Image-to-Video: which Wan 2.7 model to use

Both cost the same per second and share the same 720p/1080p tiers, so the choice comes down to what you're trying to hold constant. Image-to-Video treats your upload as the literal starting (and optionally ending) frame of one shot -- pick it for a single planned composition or transformation. Reference-to-Video treats your upload(s) as an identity anchor across many separate generations, capped at 10 seconds each -- pick it when the same character, product or style needs to recur across a batch of clips with different actions.

What creators use Wan AI 2.7 Reference-to-Video for

  • Keep a brand mascot identical across a batch of 10-second ads by reusing the same reference set for every generation.
  • Turn a character sheet into a series of 1080p story beats where the face, outfit and proportions hold from clip to clip.
  • Guide a new shot with an existing video reference when a look is easier to show than to describe in a prompt.

How to use Wan AI 2.7 Reference-to-Video on VdoBloom

  1. 1Open the Image to Video tool and select Wan AI 2.7 Reference-to-Video from the model picker.
  2. 2Upload your starting image and describe the motion you want in the prompt.
  3. 3Choose your duration (5–10s), quality (720p/1080p), review the credit cost, and generate.

Frequently asked questions

How much does Wan 2.7 Reference-to-Video cost?

Roughly 6.4 credits/second at 720p, 9.5-9.6 credits/second at 1080p. A 5-second 720p clip is 32 credits, 10 seconds is 64; at 1080p those are 48 and 95 credits.

How is Reference-to-Video different from Wan's regular image-to-video?

In standard image-to-video, your picture becomes the literal first frame of the clip. Here the uploaded references guide the subject's identity and style while your text prompt drives the action, so the same character or product stays recognizable across many separate generations rather than one shot.

What durations and resolutions does it output?

5 or 10-second clips at 720p or 1080p -- shorter than Wan 2.7 Image-to-Video's 5-15 second range, since R2V is built for consistency across multiple generations rather than one longer shot.

Can I use a video as the reference instead of images?

Yes. The model accepts reference images or a reference video, so an existing clip can serve as the guide for subject and style when stills alone don't capture what you're after.

Is Wan 2.7 open source?

No -- Alibaba hasn't published Wan 2.7's weights, unlike the earlier Wan 2.1/2.2 generations. VdoBloom's hosted access, billed in credits, is how you run it without a local install.

How strict is Wan 2.7 Reference-to-Video's content filter?

It carries VdoBloom's HOT label, meaning more permissive, flexible filtering than typical mainstream video models. Standard platform rules still apply, but creative prompts that stricter models decline are less likely to be blocked here.

Why is R2V limited to 10 seconds when Image-to-Video goes to 15?

VdoBloom's spec sheet caps R2V at 5 or 10 seconds versus Image-to-Video's 5/10/15-second range -- the shorter ceiling is a fixed limit of this variant, not a setting you can override. If you need a longer single shot without reference-locking, Wan 2.7 Image-to-Video or Wan 3.0 (up to 30 seconds) are the options.

What makes a good reference image for R2V?

A clear, well-lit image of the subject with the identifying details you want preserved -- face, outfit, product design -- visible and unobstructed. Since the prompt controls the action while the reference controls identity, keep the reference focused on the subject and let the text describe what happens.

Related models