Guides7 min readSeptember 20, 2026

30 Genjutsu Labels That Swap the Right Thing

30 tested Genjutsu labels for object swap and motion transfer, plus why vague labels like 'the player' swap the wrong person in your clip.

A Genjutsu label works when it identifies one thing and nothing else — a named person, a specific item, or a described place — because the label is the only signal the model gets for what to change; there is no mask, no bounding box, and no click-to-select region underneath it. Vague labels like “the player” or “the car” leave the model guessing, and in a scene with more than one player or car, it guesses wrong.

Disclosure: VdoBloom is our platform. Examples below come from our own generations on 20 September 2026.

Why specificity is the whole mechanism

On Genjutsu, you upload a source video plus up to 10 reference photos, and each photo gets a label field. You never see the actual prompt sent to the model — the server builds it from your labels. For an Object Swap job, your labels turn into something like: “@Video1 is the source footage. Keep every movement, the camera motion, the framing and the timing exactly as in @Video1. Change ONLY the following, and nothing else: @Image1 replaces the man on the left; @Image2 replaces the red car. Match each reference photo’s look, material and colour. Do not change the clothing or appearance of anyone not listed above.” For Motion Transfer, it’s the mirror image: everyone and everything about the people stays locked, and only the location named in your note and photos gets replaced.

That instruction is generated mechanically from your label text. If your label is “the man on the left,” that phrase is what tells the model which pixels to touch. If two people could fit that description, the model has no other way to disambiguate — it isn’t looking at a segmentation mask, it’s reading a sentence. Any photo you don’t label at all becomes “additional visual reference,” which is useful for style but does nothing to pin down a target.

Object Swap labels: people & clothing

  • the man in the blue hoodie on the left
  • the woman in the white dress closest to the camera
  • the striker in the number 10 shirt
  • the goalkeeper’s kit
  • the man’s jacket (not his face)
  • the bride’s veil
  • the runner in lane 4
  • the child in the red coat

Notice the pattern: a position (“on the left,” “closest to the camera”), a number, or a distinctive colour. Any of those three is usually enough to disambiguate one person from a crowd.

Object Swap labels: products, props & packaging

  • the coffee cup on the table
  • the phone case in her hand
  • the bottle on the counter
  • the backpack on his shoulder
  • the watch on his left wrist
  • the guitar leaning against the wall

Object Swap labels: signs, text & screens

  • the neon sign above the door
  • the storefront sign
  • the billboard in the background
  • the laptop screen

Object Swap labels: vehicles

  • the red car
  • the delivery van parked on the street
  • the motorcycle in the foreground
  • the taxi in the middle of the frame

Motion Transfer labels: locations

Motion Transfer keeps every person as they are and only swaps the environment, so your labels here describe places, not people:

  • the street
  • the sky at dusk
  • the beach at sunset
  • the rooftop skyline
  • the alley wall

These work best combined — two or three wide photos of the same place labelled individually (one for the street, one for the sky, one for a background wall), rather than one photo trying to cover the whole scene.

What to put in the note, not the label

The optional note (up to 500 characters) is where you describe the overall change and anything that must stay untouched — labels are for pointing at individual photos, the note is for instructions that apply to the whole edit:

  • “keep her hair as it is”
  • “a neon Tokyo street at night”
  • “match the warm evening light already in the shot”

Keep labels themselves under 120 characters and pointed at one thing. If you need to protect something from being changed, say so in the note rather than trying to cram it into a label.

Label → target → photo: quick reference

LabelWhat it targetsPhoto to supply
the man in the blue hoodie on the leftone specific person, Object Swapclear front-facing photo of the replacement person
the striker’s kitclothing only, not the player’s face or bodyflat-lay photo of the kit alone
the red carone specific vehicle, Object Swapside or three-quarter photo of the replacement car
the neon sign above the doora sign or piece of textphoto of the new sign with legible text
the coffee cup on the tablea small propproduct-style photo, plain background
the street / the sky at duskbackground and environment, Motion Transfer2–3 wide photos of the target location

Labels that fail, and why

  • Vague nouns: “the player,” “the car,” “the person.” If more than one thing in the source video matches the noun, the model has to pick one, and it doesn’t always pick the one you meant.
  • Two targets crammed into one label: “the man and the car” against a single reference photo confuses the model about which part of the photo maps to which part of the video. Use two photos with two separate labels instead.
  • Describing the photo instead of the target: a label like “a man in a leather jacket” describes what’s in your reference image, not which person in the source video to replace. Labels should point at the video, not narrate the photo.
  • Asking for motion changes: Genjutsu keeps the original movement, camera motion, framing and timing on purpose. A label or note asking for “a slower walk” or “turn to face the camera” won’t do anything — motion is locked by design, not by your prompt.

What we saw testing this on 20 September 2026

We ran a football clip through Object Swap with the label “the striker’s kit” and a flat-lay photo of a long black leather coat, trousers and boots. The swap was clean: the kit became the coat, the player’s face and movement didn’t change. In a second test we named two players (“the striker in the number 10 shirt” and “the defender in the number 2 shirt”) against the same coat photo, and both swapped correctly — but an unnamed third player visible in the wide shot was also altered, because nothing in our labels told the model to leave everyone else alone. The lesson: specific labels beat vague ones, but they only protect what they explicitly name. Anyone or anything you don’t label is fair game unless your note says otherwise. For Motion Transfer, the cleanest results came from wide photos of the place plus a short note (“a neon Tokyo street at night”) rather than one photo asked to carry the whole environment.

Before you submit the job

A few limits worth knowing before you upload:

  • Up to 10 reference photos per job, each with its own label.
  • Source clips run 2–30 seconds at 480p–720p; output matches the source length and aspect ratio at 480p, 720p, or 1080p.
  • Generation typically takes 4–7 minutes.
  • A 15-second clip costs 212 credits at 480p, 473 at 720p, or 852 at 1080p — draft at 480p first, then re-run the labels that worked at full resolution.

If your job needs to swap an object without touching motion, start from AI Object Swap Video. If you’re replacing the location instead of a person or item, AI Motion Transfer Video is the same engine tuned for that case. Both live inside Genjutsu AI Video, and if you want the underlying mechanics explained end to end, see Genjutsu AI, Explained.

Ready to try it?

Create your first AI video in minutes — no credit card required.

Start Creating Free →