How to Change an Outfit in a Video With AI
Swap a person's outfit in a real video clip with AI: presets, per-second pricing, and a 15-second football-kit test at 212-852 credits.
To change an outfit in a video with AI, upload the clip and a labelled reference photo of the new clothing to a video-editing model that keeps the same face, pose, camera motion and clip length — VdoBloom’s Genjutsu tool does this on ByteDance Seedance 2.5’s edit mode for roughly 41 to 912 credits depending on clip length and output resolution.
Disclosure: VdoBloom is our platform. Figures below are from our own generations on 20 September 2026.
Why this is harder than swapping an outfit in a photo
An outfit swap on a still image only has to look right in one frame. A video has to hold that outfit steady while the person moves, turns, and gets closer to or further from the camera — without the face drifting, the walk changing, or the background flickering. That is a video-editing problem, not an image-generation problem, which is why Genjutsu runs on Seedance 2.5’s edit mode rather than a text-to-video model: it treats your clip as the fixed skeleton (length, aspect ratio, faces, poses, camera motion, timing) and only repaints the region you label.
Genjutsu’s two presets
Genjutsu, at /dashboard/video-creation/genjutsu/, gives you two ways to use that edit mode:
- Object Swap — changes only what you label. Good for a wardrobe change: label the jacket, the dress, the shoes, and nothing else in the frame moves.
- Motion Transfer — keeps the people and their motion, but replaces the location around them. Not what you want for an outfit change, but useful if you are also relocating the scene.
For an AI wardrobe change, Object Swap is the preset to pick.
What you need before you start
- A source clip 2–30 seconds long, 480p–720p (854×480 to 1280×720), 24–60 fps, under 200 MB, in MP4 or MOV.
- Up to 10 reference photos of the new outfit, each with a short text label describing exactly what it replaces — for example, “the woman’s red dress,” not just “the dress.”
- An optional one-line note for anything that must stay put, such as “keep her hair as it is.”
- A choice of output resolution: 480p, 720p, or 1080p.
The price is shown from the clip length the moment you pick the file, before anything uploads, so you know the cost before you commit. Clips outside the spec above are rejected before you are charged, and a job that fails is refunded automatically.
What it costs
Genjutsu bills per second of clip at your chosen output resolution. The cheapest possible job — a 4-second clip at 480p — is 41 credits. Here is how it scales:
| Clip length | 480p | 720p | 1080p |
|---|---|---|---|
| 4 seconds | 41 credits | 92 credits | 165 credits |
| 15 seconds | 212 credits | 473 credits | 852 credits |
| 30 seconds | 408 credits | 912 credits | 1,644 credits |
New accounts get 10 free credits, which will not cover a real outfit swap — it is enough to see the interface, not to run a job. One-time credit packs start at $2.49 and never expire, and there is no plan or subscription gate on Genjutsu.
Our test: swapping a football kit mid-action
We ran a 15-second 1280×720 clip of a footballer dribbling, doing a knee-slide celebration, and a close-up. We uploaded one flat-lay photo of a long black leather coat, dark trousers, and brown boots, labelled “the striker’s kit,” and rendered at 480p.
The result: the coat followed him through the dribble, the knee-slide, and the close-up. His face, stride, and the ball itself were unchanged. In the frames where he was not the focus, the footage was near-identical to the source. The job took about four minutes, inside the tool’s 4–7 minute range.
Where it gets harder: more than one person
We also tried a version naming two players — the number 10 striker and the number 2 defender — both pointed at the same coat photo. Genjutsu dressed both of the named players correctly — but in a wide shot, it also touched an un-named third player who happened to be standing nearby. This is the honest limit: crowded scenes with multiple people increase the chance the model touches someone you did not intend to change.
What reduces this:
- Use one distinct reference photo per person you are dressing, rather than reusing the same photo for two people.
- Write specific labels — “the striker in the number 10 shirt,” not “the player” — so the model has something to key on besides position in frame.
- Use the note field to say what must not change, e.g. “do not change anyone else’s clothing.”
- Draft at 480p first. It is cheaper and fast enough to check whether the labels are landing on the right person before you pay for a 1080p render.
Step by step
- Open Genjutsu and choose the Object Swap preset.
- Upload your source clip. The credit cost appears immediately, before upload finishes.
- Upload one reference photo per item or person you want to change, and label each one specifically.
- Add a one-line note for anything that must stay the same, if needed.
- Pick your output resolution — 480p to check the result cheaply, 720p or 1080p once you are confident.
- Generate and wait roughly 4–7 minutes.
Choosing an output resolution
Output resolution is a separate choice from your source resolution, and it drives most of the cost. A 1280×720 source clip can still be rendered out at 480p, 720p, or 1080p — the model does not have to match your source. In practice, 480p is the right choice for checking whether a label is landing on the right person or item, since it costs less than half of the 720p price for the same clip. Once the labelling is confirmed, re-run the same clip and photos at 720p or 1080p for the version you actually publish. This two-pass approach costs more in total credits than guessing once, but it costs far less than paying for a 1080p render that turns out to have dressed the wrong person.
Common questions
- Can I change more than one item at once? Yes — upload up to 10 reference photos in the same job, each with its own label, so you can swap a jacket, trousers, and shoes together in one render.
- Does it work on a clip with camera movement or zoom? Yes, camera motion is one of the things the edit mode is built to preserve, along with the person’s pose and timing — that is what distinguishes it from a text-to-video regeneration.
- What file formats and sizes are accepted? MP4 or MOV, 2–30 seconds, 854×480 to 1280×720 source resolution, 24–60 fps, under 200 MB.
- What happens if my clip does not meet the spec? It is rejected before you are charged, and the price shown next to your file updates the moment you pick a different clip.
- What if the render comes out wrong? A failed job is refunded automatically; a job that completes but mislabels something (as in our two-player test) is not a technical failure, so re-running it with tighter labels is the fix, not a refund request.
Related reading
For the single-item case — a bag, a sign, a piece of furniture — see AI object swap for video. For how the editing mode preserves motion and where we have seen it fail, see Genjutsu explained, and the Genjutsu feature page for the current preset list and limits.
Bottom line
For a single person in a clean shot, AI outfit swapping on video now works close to one-shot: the coat-swap test held up frame to frame with no visible drift in face or motion. For scenes with several people close together, budget for a cheap 480p draft first and expect to tighten your labels — the model can still misattribute an un-named person in a crowd, and pretending otherwise would be the kind of hype this guide is trying to avoid.
Ready to try it?
Create your first AI video in minutes — no credit card required.
Start Creating Free →