Guides8 min readSeptember 20, 2026

AI Video Localization: Change the Background, Not the Ad

Localize a video ad for a new market by changing only the background with AI: a real stadium-to-port test, per-market cost, and a workflow table.

To localize a video ad for a new market with AI, you do not need a reshoot: VdoBloom Genjutsu’s Motion Transfer preset keeps every person, face, garment, camera move and beat of timing in your existing clip and replaces only the location, using up to 10 reference photos of the new place plus an optional text line. A 15-second clip runs 212–852 credits depending on output resolution and takes about 4–7 minutes; it does not translate speech — that is a separate Translate & Dub step.

Disclosure: VdoBloom is our platform. Figures below are from our own generations on 20 September 2026.

What “AI video localization” actually means here

When teams say they want to “localize a video ad for different markets” or put the “same ad in a different location,” they usually mean one of two things: change the visible setting (city skyline, storefront, street signage, background scenery) so the ad reads as local, or change the spoken language and captions. These are different problems with different tools. Genjutsu’s Motion Transfer preset solves the first one — it is a video-editing model, not a translation model. It takes your original footage and re-renders the environment around the same performance. Nobody re-acts the scene, nobody re-shoots, and none of the timing changes.

If your ad also needs different spoken language or lip-synced dubbing for each market, that is VdoBloom’s Translate & Dub tab, a separate step you can run before or after Motion Transfer. This guide covers the background-swap half of the job, which is usually the more expensive half to solve with a traditional production, and the part people mean when they ask “can AI change my video’s background to another city.”

How Motion Transfer works

Genjutsu lives at /dashboard/video-creation/genjutsu/. You upload a source clip (2–30 seconds, 480p–720p, 24–60 fps, under 200 MB) and choose Motion Transfer. Instead of describing the whole new scene from scratch, you feed it reference photos of the destination — up to 10 — and label each one with the part of the frame it should inform: “the street,” “the sky at dusk,” “the building facade.” You can add one line of text on top, like “a neon Tokyo street at night,” to steer mood or lighting the photos do not cover. The model underneath is ByteDance Seedance 2.5 running in its video-editing mode, which is why the output keeps the same length and aspect ratio as the source clip — it edits the existing video rather than generating a fresh one from a text prompt.

Everything that is not the environment — people, faces, clothing, props, camera movement, the pacing of the action — is meant to pass through unchanged. That is the entire value proposition: one ad, shot once, redressed for as many markets as you have reference photography for.

What stays the same, what changes

  • Stays identical: every person in frame, their faces, their clothing, their movement and choreography, the camera’s pans/zooms/handheld shake, the overall timing and beat of the clip.
  • Changes: the location itself — ground surface, background structures, sky, signage context, ambient set dressing — guided by your reference photos and optional text line.
  • Changes only if you run it: a specific product or sign in frame, via the separate Object Swap preset, run as a second job on the Motion Transfer output.

Our test: a stadium becomes a container port

On 20 September 2026 we ran a 15-second football clip — a full celebration sequence, eleven players plus a referee, a stadium crowd in the stands — through Motion Transfer with a single reference photo of a container port and a one-line note describing it.

The result kept every player’s pose, every face, the referee, the ball, and even the boom microphone visible at the edge of the close-up. What changed was the environment: the pitch became concrete dockside paving, the stands became stacked shipping containers and gantry cranes, and the celebration that originally happened in front of the stands instead happened in front of a docked cargo ship. Nothing about the actors’ performance signalled it had been altered.

Two caveats from that run and others like it:

  • One photo worked, but two or three wide shots of the same location from different angles give the model more to work with and produce steadier, more coherent backgrounds across the full clip — especially on clips with camera movement.
  • Crowded scenes (like our stadium shot) are the hardest case for the follow-up Object Swap preset, not for Motion Transfer itself. Swapping a specific sign or product in a busy frame is more failure-prone than swapping the general environment.

A market-by-market workflow

Here is the practical shape of a localization job across several markets, using our test clip as the baseline:

MarketPhotos to gatherJobs to runApprox. cost (15s, 720p)
Original (stadium)———
Port city market2–3 wide shots of a container port/dock1 Motion Transfer473 credits
Desert market2–3 wide shots of a desert stadium exterior or open desert1 Motion Transfer473 credits
City skyline market2–3 shots of the target skyline + one local sign/billboard photo1 Motion Transfer + 1 Object Swap473 + 473 credits

The price is charged per second at your chosen output resolution, and it is shown up front from the clip length before you even upload — so you know the cost of localizing to a fourth or fifth market before you commit credits to it. As a reference for scale: a 15-second clip runs 212 credits at 480p, 473 at 720p, or 852 at 1080p; a 30-second clip at 720p runs 912 credits. Failed jobs are refunded automatically, and one-time credit packs start at $2.49 and never expire, so there is no subscription gate forcing you into a plan just to test one market.

Localizing on-screen text and products with Object Swap

Motion Transfer changes the environment; it is not built to selectively swap a single sign, storefront name, or product package while leaving the rest of the frame untouched. For that, run Object Swap as a second job on the Motion Transfer output. It is a distinct preset for changing a specific product or sign in the already-localized clip — useful when the same ad needs a different bottle label, storefront name, or local billboard text per market.

The practical guidance from testing: Object Swap works most reliably on clean, uncluttered frames where the target object is clearly isolated. A wide crowd shot is the harder end of that spectrum — busy backgrounds with lots of overlapping elements give the model more ambiguity about what exactly to replace. If your ad has a clear single-shot product moment, that is a much easier Object Swap candidate.

What this does not do: language and dialogue

It is worth being direct about the boundary here: Motion Transfer and Object Swap operate on the visual scene. Neither one translates speech, changes the language of dialogue, or re-syncs lips to new words. If your localization plan for a market also requires the actors to appear to speak that market’s language, that is handled by VdoBloom’s separate Translate & Dub tab at /dashboard/video-creation/translate-dub/. The two tools are meant to be used together for a full localization pass: run Motion Transfer to change where the ad appears to be shot, then run Translate & Dub to change what the ad appears to say. They are independent jobs with independent costs, and you can run them in either order depending on which one you want to review first.

Step-by-step: localizing one ad for N markets

  • Pick your source clip: 2–30 seconds, 480p–720p, 24–60 fps, under 200 MB.
  • For each target market, gather 2–3 wide reference photos of the location, each covering a distinct part of the frame (ground, sky/background, structures).
  • Label each photo with what it shows (“the street,” “the sky,” “the building line”) and add one optional descriptive line if the photos do not capture mood or time of day.
  • Run Motion Transfer at your target output resolution; check the credit cost shown before upload.
  • If a specific product, label, or sign needs to change per market, run Object Swap on the result as a second job, favouring clean single-subject frames over crowded ones.
  • If dialogue also needs to change per market, run Translate & Dub separately.
  • Repeat per market — each market is its own job and its own reference photo set, so markets can be produced and reviewed independently rather than in one giant batch.

Honest limitations

  • Motion Transfer changes environment, not language — plan for Translate & Dub as a separate line item if dialogue needs to localize too.
  • Object Swap is noticeably harder to get clean on crowded frames; test on your busiest shot before assuming it will work across the whole ad.
  • Output resolution drives cost directly, so decide your delivery resolution per market before running the job rather than upscaling after the fact.
  • Source footage has a ceiling of 720p in and 200 MB, so a 4K master needs to be exported at 720p for this workflow.

Related reading

Ready to try it?

Create your first AI video in minutes — no credit card required.

Start Creating Free →