Image to Video Best Practices: Getting Clean Motion From a Still Photo
A repo-grounded checklist for animating a still photo on VdoBloom: what makes a good source image, how to write a preserve-plus-move prompt, real per-model credit costs, and how to fix warped faces and melting hands.
Clean motion from a still photo comes from three things: a photo the model can actually read, a prompt that tells the model what to preserve as well as what to move, and a duration short enough that the model never has to invent what it cannot see. Almost every warped face, melting hand or drifting background in an image-to-video clip traces back to one of those three. This is the checklist we use at VdoBloom, written against how the platform actually behaves rather than generic advice.
Start with a photo the model can read
Image-to-video models do not enhance. They extrapolate. Whatever is ambiguous in your still stays ambiguous in the video, and then it moves. A soft, low-light phone shot of a face at the edge of frame will produce a soft, low-light face that changes shape as it turns.
What works, in order of how much difference it makes:
- Subject fills a decent part of the frame. A head-and-shoulders or waist-up crop gives the model enough facial detail to hold identity across 5 seconds. A full-body shot where the face is 40 pixels tall will not hold.
- Even, directional light. Harsh mixed lighting creates shadow edges the model reads as geometry, and geometry that is not really there tends to wobble.
- One clear subject. Crowds, mirrors and busy backgrounds give the model extra faces to animate. It will animate them.
- Limbs and hands inside the frame. A hand cropped at the wrist has to be invented the moment the arm moves. Hands invented mid-motion are the classic failure.
- Sharp, not upscaled-to-death. Over-sharpened or heavily AI-upscaled inputs carry halo artefacts that become shimmer in motion.
VdoBloom accepts JPEG, PNG, WebP and GIF uploads. If your source is a screenshot or a heavily compressed social download, re-export it at full size before uploading — the compression blocks are real detail to the model.
If the photo shows a real, identifiable person, you must have that person’s consent before animating them. That is not a technicality: an animated likeness is the person’s likeness.
Pick the shortest duration that tells the story
Every extra second is more time for the model to drift away from your source frame, and it is billed. A 5-second clip that holds identity beats a 10-second clip that dissolves at second seven. Generate 5 seconds first, confirm the look, then re-run longer only if the motion genuinely needs the room.
Here is what a single image-to-video generation actually costs on VdoBloom. Credits are read straight from the platform pricing table, and the dollar figure uses the Lite plan rate where $15 buys 300 credits, so 1 credit is $0.05.
| Model (image to video) | Tier | Credits | Cost at $0.05/credit | Best for |
|---|---|---|---|---|
| Runway Gen-4 Turbo | 5s | 20 | $1.00 | Cheapest clean test pass |
| Runway Gen-4 Turbo | 10s | 40 | $2.00 | Longer takes on a budget |
| WAN 2.7 Image to Video | 720p / 5s | 32 | $1.60 | Everyday photo motion |
| WAN 2.7 Image to Video | 1080p / 5s | 48 | $2.40 | Delivery-quality single shot |
| WAN 2.7 Image to Video | 1080p / 15s | 143 | $7.15 | Long single take |
| Kling V2.5 Turbo | 5s | 40 | $2.00 | Human motion and dance |
| Kling V2.5 Turbo | 10s | 80 | $4.00 | Full dance or walk cycle |
| Seedance 2 Fast | 720p / 5s | 50 | $2.50 | Stylised, energetic motion |
| Seedance 2 | 1080p / 5s | 203 | $10.15 | Hero shot, final render only |
The practical pattern: test on the cheap row, deliver on the expensive one. A new account gets 10 free credits with no card, which is enough to see the workflow end to end before you commit — see pricing for the plan and credit-pack tiers, including packs from $2.49 that never expire.
Write the prompt as preserve plus move
The single biggest prompt mistake is describing only the motion. The model then treats everything else as negotiable. VdoBloom’s own effect templates are built the opposite way, and you can copy the shape:
- Name the action in one plain clause. She turns her head toward camera and smiles. Not a paragraph.
- List what must not change. Preserve facial features, body proportions, clothing, pose, lighting, shadows and background exactly as in the original image. This sentence does more work than any adjective.
- Constrain the cast. Only the person from the uploaded image may appear — no additional people. Without this, models happily add a second figure.
- Give the camera one instruction. A slow push-in, a locked-off frame, a gentle orbit. One move, not three.
- Close with the quality target. Photorealistic, smooth motion, clean blending.
Keep it to roughly four sentences. Long prompts dilute; each extra clause competes with the preserve instruction for the model’s attention.
Match the frame to the platform before you generate
Crop the source photo to the aspect ratio you are going to publish in, and set the generation to that same ratio. A 16:9 source forced into a 9:16 render makes the model invent the top and bottom of the scene — which is exactly the region where it drifts first. Vertical subject, vertical crop, vertical output.
When the result is wrong, change one thing
- Face changes shape: crop tighter on the subject and re-run. Then shorten the duration.
- Hands melt: choose a source photo where hands are fully visible and still, and avoid prompting hand-led actions.
- Background crawls: add the explicit preserve-background clause, and pick a simpler backdrop.
- Motion is too slow or nothing happens: the prompt is too cautious. Name a concrete physical action instead of an adjective like dynamic.
- Extra people appear: add the only-this-person clause.
- Clip looks fine then falls apart at the end: that is duration. Generate 5 seconds and cut.
Change one variable per re-run. Changing prompt, model and duration together tells you nothing about which one fixed it.
Use the effect tabs when the motion is a known one
If you want a standard photo-motion beat, you do not need to write a prompt at all. The effect tabs ship tuned templates: hair flip offers Dramatic, Elegant and Dynamic variants, and fashion walk offers Classic Runway, Street Style, Evening Gown and Casual Chic. Each template is a full preserve-plus-move prompt already written and tested, and you can still edit it before generating. For product stills rather than people, the 360 product video tool covers the turntable case.
Two habits that save credits
First, a failed generation is refunded automatically — you are not charged for a job that errors out. A generation that produces a bad-looking but technically complete video is not a failure, though, so cheap test passes are still the right discipline. Second, watermark-free downloads and commercial rights come with paid plans; if you are producing for a client, get that sorted before you render the final take rather than after.
Frequently asked questions
What resolution should my source photo be?
Big enough that the subject’s face is clearly detailed at full size — roughly 1024px on the short edge is a comfortable floor. Beyond about 2K there is no further gain, because the model works at its own render resolution.
Why does my subject change identity partway through the clip?
Usually duration, then crop. Models hold identity best in the first few seconds. Re-run at 5 seconds with a tighter crop and an explicit preserve-facial-features clause before trying anything else.
Can I animate a photo of someone else?
Only with that person’s consent. Animating an identifiable person without permission is not something to do, regardless of what a tool technically allows.
Which model should I default to for photo motion?
Runway Gen-4 Turbo at 20 credits for 5 seconds is the cheapest way to check whether your photo and prompt work. Once they do, WAN 2.7 at 1080p for 48 credits or Kling V2.5 Turbo at 40 credits are the usual delivery choices for human subjects.
Does a blurry photo get sharper in the video?
No. Image-to-video models extrapolate motion from what is there; they do not restore detail. Fix the still first.
Ready to try it?
Create your first AI video in minutes — no credit card required.
Start Creating Free →