Guides9 min readSeptember 20, 2026

Genjutsu AI, Tested: How Object Swap and Motion Transfer Work

What Genjutsu keeps, what it changes, where it slips and what it costs — from four real runs on Seedance 2.5 with frame-by-frame comparisons.

Genjutsu takes a clip you already have and gives it back with something changed: a person, an outfit, a product — or the entire location — while every movement, the camera path, the timing and everyone you did not name stay exactly as filmed. Higgsfield shipped it on 1 September 2026; VdoBloom's version, at /dashboard/video-creation/genjutsu/, runs on the same ByteDance Seedance 2.5 editing mode. This post is what we learned building and testing it, with the actual outputs, including the one that went wrong.

Disclosure: VdoBloom is our platform. Everything below comes from our own generations on 20 September 2026; Higgsfield figures are read from higgsfield.ai on the same day.

What "Genjutsu" actually is

Strip the name off and it is video-to-video editing with the performance locked. Seedance 2.5 has two ways to use a reference video. In its generation mode the clip is a loose guide — the model keeps the story beats but re-invents framing and faces. In its editing mode the clip is the ground truth: the output inherits its length and aspect ratio, and only what the instruction names changes. Genjutsu, on both platforms, is that editing mode with two preset instructions:

  • Object Swap — "keep everything; change only X to what this photo shows."
  • Motion Transfer — "keep every person and every movement; replace only the location with what these photos show."

That is the whole trick. The engineering is in how the instruction is built and in what you refuse to send the model.

Test 1 — Object Swap: a footballer's kit becomes a leather coat

Source: a 15-second, 1280×720 football clip — a dribble, a shot, a keeper dive, a knee-slide celebration, a close-up scream. Reference: one flat-lay photo of a long black leather coat, dark trousers and brown boots. Instruction: change only the clothing of the player in the number 10 shirt.

Result, frame by frame against the source at 1, 4, 7, 10 and 13 seconds: the wide shot is pixel-matched — every player, ad board and section of crowd where it was. At 4 seconds the striker dribbles past the defender in the coat, belt and all; the ball, the defender, the keeper and the goal are untouched. The keeper-dive frame, where the striker is out of shot, is near-identical to the source. The knee-slide and the scream keep his face, his expression and his pose. Output: 15.04 seconds, same ratio, 250 seconds of generation.

What we did not expect was how literally "only" is honoured. Nothing else in the frame moved.

Test 2 — Motion Transfer done wrong, then right

Our first Motion Transfer run used the generation mode with the port photo as a scene reference. It produced a good video — containers, cranes, the sequence of kick, celebration, huddle in the right order — but the framing changed at every timestamp and the players had new faces. Story-level transfer, not motion transfer.

The second run used the editing mode with the instruction "keep every person exactly as they are; change only the location to the container port in the photo." That one is the Motion Transfer card on the tab: all eleven players, the referee, the ball, the poses and even the boom microphone in the close-up are identical to the source; the pitch became concrete, the stands became container stacks and cranes, the scream now happens in front of a ship. Same faces. That difference — mode, not model — is the single most useful thing we learned, and it is why the tab never offers the generation path.

Test 3 — two people at once, and where it slipped

We asked for both the number 10 striker and the number 2 defender to be dressed in the same coat. Both were, with faces and motion kept. But in the wide shot at least one un-named player also got a coat. Two lessons went straight into the product: the instruction now ends with an explicit "do not change anyone not listed", and the interface makes you label each reference photo with what it replaces — "the striker in the number 10 shirt", not "the player" — because specificity is what the model keys on. Crowded scenes are still the hardest case; we say so on the tab rather than promising pixel-precise targeting.

What the model refuses, and why we check first

Seedance's editing mode has hard input rules: 2-30 seconds, 480p-720p sources (roughly 854×480 to 1280×720), 24-60 fps, under 200 MB. It also insists on two API parameters — duration −1 and an adaptive ratio — and rejects the job otherwise with a message saying it "identified your task as video editing". None of this is in the public documentation; we found it by being rejected. The consequence for a product is that a clip must be probed before credits are charged, because the model's own rejection arrives after. VdoBloom reads the clip's length from the file the moment you choose it, shows the price, and refuses out-of-spec clips before anything moves. One thing we got wrong in the first build and fixed the same day: a reused helper capped clip length at 15 seconds, which would have under-billed every long clip — worth mentioning because it is exactly the class of bug this feature invites.

What it costs, honestly

Editing mode is billed on the with-reference-video rate and the output is as long as the input, so the price is a function of clip length and output resolution only. On VdoBloom a 15-second clip is 212 credits at 480p, 473 at 720p and 852 at 1080p; 30 seconds at 720p is 912. Higgsfield lists roughly 40 / 104 / 144 credits for the same 15 seconds at 480p / 720p / 1080p (about $2.00 / $5.20 / $7.20), but Genjutsu there requires Seedance 2.5, which its $9 Basic plan excludes, so it effectively starts at the Pro plan and the credits expire monthly. VdoBloom has no plan gate and its one-time packs, from $2.49, never expire. Per clip Higgsfield is cheaper; per occasional use VdoBloom is. We would rather you knew that than found out.

When to use which preset

You wantPresetGive it
A different outfit, product, prop or person in the same shotObject SwapOne photo per change, each labelled with what it replaces
The same people and action somewhere elseMotion TransferPhotos of the new place, labelled by what they show
BothRun Motion Transfer, then Object Swap on the resultTwo jobs; each keeps what the other changed

What we would tell a friend

Label everything. Use distinct photos for distinct targets. Draft at 480p (a 15-second draft is 212 credits) and only re-run the keeper at 720p or 1080p. Expect four to seven minutes. And if a result changes something you did not name, name it in the note and run again — the model is obedient; it just needs to be told what "only" means.

Genjutsu is at /dashboard/video-creation/genjutsu/. The feature page is Genjutsu AI video; the head-to-head with Higgsfield is here.

Ready to try it?

Create your first AI video in minutes — no credit card required.

Start Creating Free →