Guide7 min readSeptember 13, 2026

How to Keep an AI Video Character Consistent Across Shots on VdoBloom

The reference-image workflow that stops AI characters drifting between shots, with real VdoBloom credit costs for every step from hero still to 1080p final.

An AI character stays consistent across shots only when you stop re-describing them in words and start re-feeding the same locked reference image into every single generation. Text prompts cannot pin an identity — every render is an independent roll of the dice, and a phrase like “a woman in her late twenties with dark curly hair” describes millions of faces. The workflow that actually holds up is: make one still you are happy with, treat it as the only source of truth, and drive every shot from that image rather than from fresh text.

Why your character drifts between shots

Video models have no memory between jobs. Shot one and shot two are two separate requests, and nothing carries over except what you upload. Three things cause most of the drift people complain about:

  • Text-only generation. A text-to-video prompt re-invents the face each time. Even with an identical prompt and a fixed seed, small sampler differences change bone structure, hairline and skin tone.
  • Different source stills. People animate three different photos of the same person, shot in three different lightings, and then blame the video model for the mismatch. The model was faithful — to three different inputs.
  • Prompts that invite change. Words like “stylish” or “glamorous” are permission slips, and your character comes back with a new jawline.

Step 1: build one locked reference still before you touch video

Do the identity work in the image tools, where a retry costs cents instead of dollars. On VdoBloom the image editors accept multiple reference images at once, which is exactly what a character sheet needs: Nano Banana Edit takes up to 10 reference images for 5 credits, Nano Banana 2 takes up to 14 for 15 credits, and Nano Banana Pro runs 18 credits at 1K or 2K and 24 at 4K with resolution control.

Generate or edit until you have one hero still — front-facing, even lighting, whole outfit visible. Then, in the same editor, produce the variations you will need as shot starters: the same character at three-quarter angle, from behind, in the second location, holding the prop. Because those variations come from the hero still as a reference image, they inherit the face instead of guessing it. Do this in the image editor and keep every output; those stills are your continuity bible.

Step 2: animate from the still, never from text alone

Once the stills exist, every video shot becomes an image-to-video job. The image carries identity, the prompt carries motion only. That split is the whole trick. A prompt for an image-to-video shot should describe what moves and how the camera behaves — nothing about who the person is, because the person is already in the frame you uploaded.

If the reference photo is of a real person, you must have that person’s explicit consent before animating them, and that applies to every shot in the sequence, not just the first.

Step 3: copy the identity-lock sentence the effect tabs already use

VdoBloom’s photo-motion effect tabs are built around a prompt pattern that exists purely to stop drift. Look at the default prompt behind a tab like Fashion Walk and you will find this instruction near the top: preserve facial features, body proportions, clothing, lighting, and identity exactly as in the original image. The two-person tabs add a second clause: only the people from the uploaded images may appear in the scene, no additional persons or characters.

Those sentences are not decoration. Paste them into your own custom prompts on the main generator too. The first blocks the model from re-imagining the face; the second stops the extra background people that quietly break continuity when a crowd appears in shot three and not shot two.

Two-image models: know what the second image actually means

This is where most multi-shot projects go wrong. Two models can both accept two images and mean completely different things by them:

  • Seedance 2, 2 Fast, 2 Mini and 2.5 take genuine multimodal reference images, tagged in the prompt as @Image1 and @Image2. Two photos means two characters in one scene.
  • Veo 3 Fast and Veo 3 Quality and the Gemini Omni family also take multiple reference images natively.
  • Kling 3.0 Omni treats a second image as an end frame. Feed it two people and the video morphs person one into person two.
  • WAN 3.0 uses first-frame and last-frame semantics for the same reason.

For two characters who must both stay themselves, stay on the Seedance 2 family or Veo 3. For a deliberate transformation between two states of one character, end-frame models are the right tool — that is a feature, not a bug.

Draft cheap, finish expensive

Consistency work is iteration work, and iteration should not cost premium credits. At 300 credits for the $15 Lite plan, one credit is $0.05, so here is what each step of the ladder really costs:

Step in the workflowModel and tierCreditsCost
Build the hero stillNano Banana Edit, up to 10 refs5$0.25
Character sheet, many refsNano Banana 2, up to 14 refs15$0.75
High-res hero stillNano Banana Pro, 4K24$1.20
Throwaway motion testSeedance 1.5 Pro, 480p, 4s3$0.15
Readable draft shotSeedance 2 Mini, 720p, 5s17$0.85
Two-character draftSeedance 2, 720p, 5s82$4.10
Final delivery shotSeedance 2, 1080p, 5s203$10.15
Alternative finalVeo 3 Fast, flat rate40$2.00
Alternative finalWAN 3.0, 720p, 5s32$1.60

A 3-credit 480p test at 15 cents tells you whether the face holds. Only when the motion and the identity both survive that test should you spend 203 credits on the 1080p version. New accounts get 10 free credits with no card, which is enough for three of those motion tests before you commit to anything. Full plan details are on the pricing page.

A worked four-shot sequence

  1. Hero still. Generate or upload one front-facing portrait. Edit until the outfit and lighting are final. Cost: 5 to 24 credits.
  2. Shot starters. In the image editor, use the hero still as a reference to produce three more stills: three-quarter angle, walking away, and a second location. Cost: 5 credits each.
  3. Motion tests. Animate all four stills at Seedance 1.5 Pro 480p 4s, 3 credits each, 12 credits total. Compare the four faces side by side. If one drifts, fix the still, not the prompt.
  4. Final renders. Re-run the approved four at 1080p, keeping the exact same stills and prompts. Change nothing else.

The discipline that matters is in step three: when a shot drifts, adding adjectives is the wrong reflex. Go back and fix the input image.

When consistency still fails

If the face holds in three shots and collapses in the fourth, check the outlier for the usual suspects: a much longer duration, a much wider camera move, a resolution change mid-sequence, or an extra person in the prompt. Long clips drift more than short ones, so build a 24-second scene as four 6-second shots rather than one long take. Wide shots give the model fewer face pixels to preserve, so keep identity-critical beats in mid or close framing. And if you inherited a look you like but cannot describe, run the still through photo to prompt and reuse the wording it gives you across every shot. If you are still choosing a base model for the sequence, the image-to-video model comparison is the faster way to narrow it down, and the Seedance prompt library shows the phrasing that survives multiple renders.

Frequently asked questions

Does a fixed seed keep my character consistent?

Not reliably, and not across different prompts or models. A seed only makes one exact request reproducible. Change a single word, the duration or the resolution and the face moves. A reference image is the only control that survives those changes.

Can I keep one character consistent across shots with text-to-video alone?

No. Generate the character once as an image, then use image-to-video for every shot. This is the single biggest quality jump available to most users, and it usually costs less than re-rolling text-to-video clips until one matches.

Which models accept two different people without merging them?

The Seedance 2 family, Veo 3 Fast and Veo 3 Quality, and the Gemini Omni models take true multi-reference inputs. Kling 3.0 Omni and WAN 3.0 treat a second image as an end or last frame, so they will morph one person into the other.

How much should I budget for a consistent four-shot sequence?

Roughly 30 credits of image work, 12 credits of 480p tests, then the finals: 160 credits at Veo 3 Fast, or 812 at Seedance 2 1080p 5s.

Do I need permission to animate a photo of someone else?

Yes. If a real person appears in the source photo, you need their consent before you animate them, and paid plans also carry commercial rights for the output, which makes that consent more important rather than less.

Ready to try it?

Create your first AI video in minutes — no credit card required.

Start Creating Free →