AI Lift and Carry Video Generator: One or Two Photos, One Cinematic Lift
The VdoBloom AI Lift and Carry Video Generator animates one person lifting and carrying another, generated from photos rather than filmed. It handles the part that is genuinely hard to shoot - two people, a lift, and a camera that has to be somewhere - and produces a short clip where the moment reads as real. Nobody has to actually pick anyone up.
There are two ways in, and the photo-count toggle at the top of the tab is what switches between them. Two-photo mode, which is where the tab opens, takes a separate portrait of each person and places both into one generated scene, matching lighting and perspective so neither looks pasted in. Single-photo mode takes one image that already contains both people, which usually tracks the real pair more closely because the model is not reconciling two different cameras.
Three templates set the emotional register, and they are meaningfully different. Playful is the romantic version, with joyful expressions and easy energy. Dramatic pushes into film territory: intense emotion, strong contrast, the lift as a beat in a story. Sweet dials it down to a gentle, tender carry under soft light. The prompt is editable underneath all three, which is how you decide who lifts whom, where they are, and how the shot moves - naming the roles outright is the single most useful edit on a two-person effect.
New accounts start with 10 credits. In two-photo mode the picker runs on the native effect roster - 17 variants across 5 families: Runway, Kling, PixVerse, Vidu and LTX-2 - where Runway Gen-3 at 720p for 5 seconds costs 10 credits, exactly what the starter balance covers. Switch the toggle to single-photo mode and the full image-to-video roster opens up, 44 variants across 15 families, including Seedance 1.5 Pro at 8 credits for the same 720p, 5-second render. Free-tier clips come back watermarked and cannot be downloaded; a paid plan removes the watermark and unlocks the file.
How it works
- 1
Decide one photo or two
Use the photo-count toggle. Two-photo mode, the tab's starting state, takes a separate portrait of each person and composites them. Single-photo mode takes one image that already contains both, which usually holds the likeness better - and it also opens a much larger model picker.
- 2
Upload the photos
Clear, well-lit shots where both faces are visible and reasonably large in frame. Full-body or three-quarter framing gives the model more to work with than tight head-and-shoulders crops, because a lift depends on it understanding relative height and where the limbs start.
- 3
Choose the mood
Playful for a warm romantic lift, Dramatic for an intense cinematic one, Sweet for a gentle tender carry. Each template writes a starting prompt, and all three are worth running once you have found a photo pairing that works.
- 4
Say who lifts whom
Edit the prompt to name the roles explicitly - which person does the lifting, which is carried - along with the setting and the camera move. 'The person in the first photo lifts the person in the second' removes guesswork the model would otherwise resolve at random.
- 5
Pick the model and settings
In two-photo mode the picker is restricted to the 17 native effect variants, and Runway Gen-3 at 720p and 5 seconds is both the default and the cheapest at 10 credits; PixVerse v3.5 and Vidu Q2 Turbo are the next step up at 18. In single-photo mode the full roster opens and Seedance 1.5 Pro drops the same render to 8. Set duration and aspect ratio before rendering - the tab opens on 16:9, so switch to 9:16 for short-form.
- 6
Generate and review
After a few minutes the clip lands in your library. Check the faces and the arms first - those are where two-person scenes go wrong. Free-tier renders are watermarked and locked to the library; a paid plan removes the watermark and enables the MP4 download.
Why use this tool
Two separate photos, one believable scene
You do not need a picture of the two people together. Upload a portrait each and the model stages them in a shared scene, matching light direction and perspective so neither reads as a cut-out dropped on top of the other.
You choose who lifts whom
The prompt controls the roles outright. Naming the person doing the lifting and the person being carried removes the guesswork that makes most two-person generators unpredictable, and it is the fastest fix when a first render gets it backwards.
Three moods that actually differ
Playful, Dramatic and Sweet are distinct prompt recipes with different lighting and energy, so a joke clip for a friend and a film-styled couple moment come from the same upload without you writing either prompt from scratch.
No per-image surcharge, but the roster changes
Price follows the model, resolution and duration, not how many photos you upload. What does change is the picker: two-photo mode runs on the 17 native effect variants where the cheapest option is Runway Gen-3 at 10 credits for 720p and 5 seconds, while single-photo mode opens the full 44-variant roster and its cheaper entries.
Two rosters behind one tab
Two-photo mode gives you 17 variants across 5 families - Runway, Kling, PixVerse, Vidu and LTX-2. Single-photo mode gives you 44 across 15. Draft on the cheap end, then re-run the prompt that worked on a heavier family when the lift needs to look like a film frame that moves.
A free render, with the terms stated
The 10 starter credits cover one Runway Gen-3 render in two-photo mode at 720p and 5 seconds. It plays back at full resolution in your library with a VdoBloom watermark and the download disabled, which is enough to judge the likeness before a paid plan clears both.
Frequently asked questions
How does the AI combine two photos into one lift video?
In two-photo mode the model reads both portraits, places the two people into a single generated scene, and animates the lift while matching lighting and perspective so they look like they were photographed together. Faces and clothing carry across from your uploads. If you already have a picture of both people, single-photo mode usually gives a closer likeness, because there is only one camera and one light source to honour.
Who ends up doing the lifting?
You decide, through the prompt. The templates supply a starting description, and editing it lets you name which person lifts and which is carried, plus the setting and the emotional tone. Being explicit - 'the person in the first photo lifts the person in the second' - gives far more consistent results than leaving the model to choose.
Is it free to make a lift and carry video?
Your first render is covered. New accounts receive 10 credits, and in two-photo mode the cheapest model available is Runway Gen-3 at 10 credits for 720p and 5 seconds - the starter balance exactly. Switch to single-photo mode and the picker widens to include Seedance 1.5 Pro at 8, which leaves you change. Free-tier clips are watermarked and cannot be downloaded until you are on a paid plan.
Does two-photo mode cost more than one photo?
There is no per-image surcharge - the credit price is set by the model, resolution and duration. What changes is which models you can pick. Two-photo mode is restricted to the native effect roster, whose cheapest option is Runway Gen-3 at 10 credits, while single-photo mode opens the full image-to-video roster where Seedance 1.5 Pro is 8 and ByteDance V1 Pro Fast I2V is 7. So in practice the two-photo route has a higher floor.
Will there be a watermark on the clip?
On the free tier, yes. Free-tier renders play in your library with a VdoBloom watermark over the video, and download, right-click saving and picture-in-picture are disabled. A paid plan removes the watermark and releases the MP4 - which is what you want before sending the clip to the person in it.
Do I have to create an account?
Yes, and it is free with no card. The account holds your 10 starter credits and is where both uploaded portraits and the finished render are stored - on a two-photo effect that matters, because you are trusting the tool with two people's pictures, not just your own, and they stay tied to your account rather than going anywhere public.
How long does the render take?
Usually a few minutes. Two-person interaction is heavier than a single-subject animation, so it tends to sit slower than a portrait clip on the same model, and longer durations and premium families add more. Generation happens server-side, so you can close the tab and collect the video from your library afterwards.
What photos work best for two-person effects?
Sharp, evenly lit images where each face is clearly visible and not too small in the frame. Full-body or three-quarter shots help the model understand posture and proportion, which a lift depends on more than a static effect does. Sunglasses, heavy shadow across a face, extreme angles and very low resolution are the most common causes of a poor likeness.
Can I upload a photo of someone else?
Only with their consent, and only if they are an adult. Treat a friend's or partner's picture as theirs to approve. Uploads and generated videos stay tied to your account and are not published by VdoBloom. Generations are moderated and requests that trip those checks can be rejected.
What is the difference between lift and carry, bridal carry, and carry me?
This tool covers general lifting moments across three moods - playful, dramatic and sweet - and leaves the roles entirely to your prompt. The bridal carry generator is tuned for the classic romantic pose and includes a golden-hour sunset variant, so it is the better choice for wedding content. Carry Me handles the comedic carries: piggyback rides and fireman lifts.