VdoBloom
← Back

Realistic Japanese Room Selfie Vlog

Vlog / Social LifestyleUGC / Talking Head Adby @タナベ | AI動画 × マーケティング · source ↗September 6, 2026

Full prompt

Read the contents of these URLs and turn them into a Skill. Always use this Skill when requesting prompt generation for Seedance 2.5. Remember this so you don't forget. https://t.co/BPHSzqNopm https://t.co/aUJgXKFRip https://t.co/u79zESgiTk Also, add the following rules to the Skill. When generating dialogue for lip-syncing, follow these rules: - Represent difficult kanji that are hard for AI to read in hiragana. - Do not represent numbers in kanji (e.g., Two Thousand Years, One Hundred and Ninety-Two times); represent numbers using Arabic numerals (e.g., 2,000 years, 192 times). - Represent English in katakana. [Format] 20 seconds, vertical 9:16. A daily selfie video taken with the front camera of a smartphone. A single continuous shot from beginning to end. The camera's left-right tilt (roll) is 0 degrees throughout. Keep the vertical screen horizontal and do not use a Dutch angle tilted diagonally. [Person/Reference] Attached Image 1 is a character sheet of the same single female character, combining a close-up of the face, a full-body front view, and a full-body back view. Use it as the sole primary reference for the person and clothing. It is not the starting frame. Do not carry over the sheet's gray background, standing posture, studio lighting, or 3-view layout into the video. Do not show the sheet itself or switch scenes from the sheet. Live-action footage starting from the first frame where the woman walks into the following room holding the camera. The woman is a fictional 29-year-old, 163 cm tall Japanese person. Maintain the facial structure, slender build and shoulder width, delicate arms, slender neck, and dark brown bob with thin bangs from Image 1 throughout. A clean oval face, naturally sized eyes, and modest daily makeup. She looks cool and clean, with a friendly natural smile. Use the vibrant cobalt blue fine-ribbed tank top and light blue washed wide-straight denim from Image 1 as they are. The tank top has a shallow square neck, a compact silhouette that naturally fits the body, and a length that reaches near the waist of the denim. The denim fits naturally around the waist and is loose from the thighs to the hem. Maintain the waist position, hem length, pockets, and fading from Image 1. Barefoot, no glasses or accessories. Maintain the slender build from Image 1 even when seated. [Action] Continuously depict the following composition and flow of movement over 20 seconds. At the start, the woman walks holding the camera near her waist. She places the camera on a desk, sits in a chair herself, and continues as if talking to one close friend. After sitting, she tucks her hair behind her ear once, and finally waves a small hand and picks up the camera. The specified conversation continues during the walking, camera setup, and sitting. [Environment] A small rental one-room apartment in Japan. White walls and ceiling, light wood grain floor. A wooden desk and chair, a large sliding window, lace curtains. On the back left of the desk is a wooden-framed tabletop mirror, a navy canvas tote bag, and a few books next to it. The front of the desk is kept clear. A normal room that is clean but has a lived-in feel. Maintain the position of objects and the shape of the room throughout. Do not reflect a legible face or a different person in the mirror. [Camera/Composition Transition] A single continuous shot moving from a low handheld position to being fixed on the desk. The character sheet is a reference for the person and clothing only, not for the camera composition. 0.0 to 4.0 seconds: Hold the vertical front camera near the waist and point the lens upward. However, only change the vertical angle (pitch) for looking up, and do not tilt the screen clockwise or counter-clockwise. In the foreground are the torso and the waist of the pants, in the upper part of the screen are the chin and face, and behind them is the wide ceiling. There is only small vertical movement matching the walking pace, no rotation to the left or right, proceeding from a dim entrance to a bright room with a window. The arm holding the camera is foreshortened at the edge of the screen, and the smartphone itself and the hand gripping it are off-screen. At the beginning, she speaks while looking in the direction she is walking. 4.0 to 5.8 seconds: Place the camera on a stable vertical stand on the wooden desk. The stand is off-screen. Lower the camera and adjust the vertical angle to continuously transition to a composition of the window, desk, mirror, and bag. Do not rotate left or right. Fingertips briefly cross the bottom edge, and small vertical shakes decay and stop. Upon completion of placement, the horizontal level is firmly set with a roll of 0 degrees and no left-right tilt. The woman temporarily moves off to the right of the screen and returns from the right to sit in the chair. Even during this short placement, the voice is heard close by. 5.8 to 19.5 seconds: The camera is fixed horizontally on a stable stand at the edge of the desk, without changing the angle or focal length. A wide vertical composition looking slightly up from a low position. The upward view here is only in the vertical angle. Rotation around the optical axis is 0 degrees; do not tilt the screen left or right. Vertical seams of the wall near the center of the screen or the vertical frame of the window are almost parallel to the vertical edges of the image. Leave a natural perspective of a light wide-angle lens, but do not make the whole room diagonal. Do not confuse the desk edge appearing diagonal due to depth with tilting the camera itself. The woman sits to the right of the screen center. She is in frame from head to waist and both forearms, leaving a wide white wall and ceiling above her head. Do not make it a close-up of the face. The large window and lace curtains are visible from behind her right side. The wooden tabletop spreads diagonally from the bottom of the screen to the back left. The wooden-framed mirror and navy bag are on the left of the screen. The woman's torso faces the mirror/desk to the left, and she turns her face toward the lens when speaking. Do not make the temporary closeness of the person seen immediately after sitting into a zoom in the fixed composition. 19.5 to 20.0 seconds: The woman reaches for the camera, her hand blurs near the lens, and she lifts it without tilting the camera left or right, ending as the screen moves slightly upward. Cuts, scene resets, zooms, panning to follow the person, Dutch angles, left-right tilting, and rotation around the optical axis are prohibited. Do not cause focal length changes during the fixed section. The background is identifiable as typical for a smartphone. Do not show the smartphone itself or other filming equipment. [Light and Presence] At the beginning, the exposure naturally changes from the dimness of the entrance to the natural light from the window. After sitting, soft natural light enters from the window at the back right of the screen along with reflections from the white walls. Maintain fine pores on the cheeks, peach fuzz, thin shadows around the eyes, fine wrinkles on the lips, and natural sebum reflections. A smile with slight asymmetry, natural blinking and breathing, and small eyebrow movements matching the content of the speech. Do not process the skin to be smooth. Do not use beauty filters, CG-like skin, or advertising-like lighting. Maintain the fine texture of smartphone footage. [Timing/Acting] 0.0 to 6.2 seconds: Line 1 starts from the first 0.2 seconds. Walk into the room, place the camera, and speak while moving to the chair. Pronounce "Seedance 2.5" clearly. Do not peer into the lens at the beginning. Use hands for camera placement and do not perform additional gestures. 6.2 to 8.8 seconds: While lowering into the chair, turn face toward the lens and deliver Line 2. A short pause before "but" (kedo). The smile fades slightly, and eyebrows are knitted slightly. Tilt head slightly as if seeking empathy during "it's difficult" (muzukashii no). Do not make a serious expression or exaggerated disappointment. 8.8 to 16.7 seconds: Line 3. Expression returns to soft. While speaking, tuck hair on one side behind the ear once and lower hand. Nod slightly at "because I wrote it in a post." Look straight into the lens for "please bookmark it," saying it in a friendly, slightly emphasizing way. Do not use pointing or large gestures. 16.7 to 18.6 seconds: Take a short breath and deliver Line 4. A smile spreads naturally, raise one hand to the side of the cheek, and wave slightly back and forth from the wrist. 18.6 to 19.5 seconds: Finish speaking and lower hand with a soft lingering smile. At 19.5 to 20.0 seconds, reach for the camera and end while picking it up. Timeframes are a guide for continuous acting. Do not stop or cut the screen at each boundary. [Audio/Dialogue] Natural Tokyo Japanese. A 29-year-old woman's calm, slightly breathy, clear voice. A quick-paced, light tempo as if talking to a friend. Speak all lines in about 18 seconds. Avoid monotone, announcer-like tone, excessively high pitch, or salesperson-like emphasis. Insert short natural pauses between sentences. Synchronize mouth movements with Japanese pronunciation. Speak only the following four sentences in order, without changing a single word. Do not add, paraphrase, or repeat. Line 1: Ai de kouiu douga o tsukuru nara "Seedance Niiten-go" ga saikou. (If you're making videos like this with AI, "Seedance 2.5" is the best.) Line 2: Kedo, puronputo ga muzukashii no. (But the prompts are difficult.) Line 3: Ii puronputo o tsukuru houhou o posuto ni kaita kara, bukku maaku shite oite ne. (I wrote how to make good prompts in a post, so please bookmark it.) Line 4: Jaa, mata ne. (Well, see you again.) "AI" is A-I, the pronunciation of "Seedance 2.5" is "Seedance Niiten-go." Read the number "Niiten-go" clearly. Do not read the quotation marks. Sound consists only of her own voice, faint breathing, very small rustling of clothes, and indoor ambient sounds. No BGM, added sound effects, narration, subtitles, on-screen text, or logos. This woman is the only character appearing. Do not deform or multiply the face, hands, clothes, or background midway.

Original prompt

これらのURLの内容を読み込み、Skillにしてください。 Seedance 2.5のプロンプト生成を依頼するときは、必ず毎回このSkillを使ってください。 忘れないように覚えておいて。 https://t.co/BPHSzqNopm https://t.co/aUJgXKFRip https://t.co/u79zESgiTk また、Skillには以下のルールも追加してください。 リップシンク用のセリフを生成するときは、以下のルールに従ってください。 ・AIが読みにくい難しい漢字は、ひらがなで表現する ・数字は漢字で表現しないで(例:二千年、一九二回)、数字は算用数字(例:2,000年、192回)で表現して ・英語はカタカナで表現する 【形式】 20秒、縦9:16。スマートフォン前面カメラで撮った日常の自撮り動画。最初から最後まで1つの連続ショット。カメラの左右の傾き(ロール)は全編0度。縦画面を水平に保ち、斜めに傾けたダッチアングルにしない。 【人物・参照】 添付の画像1は、顔のアップ・正面全身・背面全身をまとめた、同じ女性1人のキャラクターシート。人物と衣服の唯一の主参照として使う。開始フレームではない。シートのグレー背景、立ち姿、スタジオ照明、3ビューの配置を動画に引き継がない。シート自体を映したり、シートから場面を切り替えたりしない。最初のフレームから、以下の部屋へ女性がカメラを持って歩いて入る実写映像。 女性は架空の29歳、身長163cmの日本人。画像1の顔の骨格、細身の体格と肩幅、華奢な腕、すっとした首、薄い前髪のあるダークブラウンのボブを全編で維持する。すっきりした卵型の輪郭、自然な大きさの目、控えめな日常メイク。涼しげで清潔感があり、自然な微笑みに親しみがある。 画像1の鮮やかなコバルトブルーの細リブのタンクトップと、淡いライトブルーのウォッシュ加工のワイドストレートデニムをそのまま使う。タンクトップは浅いスクエアネックで、自然に体に沿うコンパクトなシルエット、丈はデニムのウエスト付近まで。デニムは腰回りが自然に合い、太ももから裾までゆったりしている。画像1のウエスト位置、裾丈、ポケットと色落ちを維持する。裸足、眼鏡とアクセサリーなし。座った姿でも画像1の細身の体格を保つ。 【動作】 20秒間で、以下の構図と動作の流れを連続して描く。冒頭は女性がカメラを腰付近に持って歩く。机にカメラを置き、自分が椅子へ座り、親しい友達1人に話すように続ける。着席後は一度だけ髪を耳にかけ、最後に小さく手を振ってカメラを取り上げる。歩行、カメラ設置、着席の間も指定の会話はつながる。 【環境】 日本の小さな賃貸ワンルーム。白い壁と天井、明るい木目の床。木の机と椅子、大きな掃き出し窓、レースカーテン。机の左奥に木枠の卓上鏡、ネイビーのキャンバストート、その横に少数の本。机の手前は空ける。清潔だが生活感のある普通の部屋。全編で物の位置と部屋の形を保つ。鏡に読める顔や別人を映さない。 【カメラ・構図の推移】 低い手持ちから机上固定へ移る1つの連続ショット。キャラクターシートは人物と衣服だけの参照であり、カメラ構図の参照ではない。 0.0〜4.0秒:縦向きの前面カメラを腰付近で持ち、レンズを上へ向ける。ただし見上げるための上下の角度(ピッチ)だけを変え、画面を時計回り・反時計回りには傾けない。手前に胴体とパンツの腰、画面の上方に顎と顔、その背後に広い天井。歩幅に合わせた小さな上下の移動だけがあり、左右への回転はなく、薄暗い入口から窓の明るい部屋へ進む。持っている腕は画面端で短縮され、スマホ本体と握る手は画面外。冒頭は進行方向を見て話す。 4.0〜5.8秒:木の机の安定した縦向きスタンドへカメラを置く。スタンドは画面外。カメラを下げ、上下の角度を調整して、窓、机、鏡、バッグの構図へ連続的に移る。左右には回転させない。指先が下端を短く横切り、小さな上下の揺れが減衰して止まる。設置完了時にはロール0度、左右の傾きがない状態で確実に水平が取れている。女性は一時的に画面右へ外れ、椅子に座るため右側から戻る。この短い設置中も声は近くで聞こえる。 5.8〜19.5秒:カメラは机の端の安定したスタンドに水平を取って固定し、角度と焦点距離を変えない。少し低い位置からわずかに見上げる、広めの縦構図。ここでの見上げは上下方向の角度だけ。光軸まわりの回転は0度、画面を左右に倒さない。画面中央付近の壁の垂直な継ぎ目や窓の縦枠が画像の縦辺にほぼ平行になる。軽い広角の自然な遠近感は残すが、部屋全体を斜めにしない。奥行きによって机の辺が斜めに見えることと、カメラ自体を傾けることを混同しない。女性は画面中央より右に座る。頭から腰と両前腕まで入り、頭上には広く白い壁と天井を残す。顔のアップにしない。大きな窓とレースカーテンは女性の右後方から背後に見える。木の天板が画面下部から左奥へ斜めに広がる。木枠の鏡とネイビーのバッグは画面左。女性の胴体は左の鏡・机の方向に向き、話しかける際には顔をレンズへ向ける。着席直後に見える一時的な人物の近さを、固定構図のズームにしない。 19.5〜20.0秒:女性がカメラへ手を伸ばし、レンズ近くで手がぼけ、カメラを左右に傾けず持ち上げ、画面が少し上へ移動したところで終了。 カット、場面リセット、ズーム、人物を追うパン、ダッチアングル、左右の傾き、光軸まわりの回転は禁止。固定区間の画角変化を起こさない。背景はスマホらしく判別できる。スマホ本体や別の撮影機材は映さない。 【光と実在感】 冒頭は入口の薄暗さから窓の自然光へ露出が自然に変わる。着席後は画面右後方の窓から入る柔らかな自然光と白い壁の反射。頬の細かな毛穴、産毛、薄い目元の陰影、唇の細い皺、自然な皮脂の反射を保つ。微かな左右差のある笑顔、自然な瞬きと呼吸、話す内容に合った小さな眉の動き。肌を滑らかに加工しない。美顔フィルター、CG的な肌、広告のような照明は使わない。微細なスマホ映像の質感を保つ。 【タイミング・演技】 0.0〜6.2秒:最初の0.2秒からセリフ1。部屋へ歩き、カメラを置き、椅子へ移動しながら話す。「シーダンス2.5」を明瞭に発音する。冒頭でレンズをのぞき込まない。手はカメラ設置に使い、追加のジェスチャーはしない。 6.2〜8.8秒:椅子へ腰を下ろしながらレンズへ顔を向け、セリフ2。「けど」の前に短い間。笑みが少しほどけ、眉をわずかに寄せる。「むずかしいの」で共感を求めるように軽く首を傾ける。深刻な表情や大げさな落胆にはしない。 8.8〜16.7秒:セリフ3。表情が柔らかく戻る。話しながら一度だけ片側の髪を耳にかけ、手を下ろす。「ポストに書いたから」で小さくうなずく。「ブックマークしておいてね」はレンズをまっすぐ見て、少しだけ念を押す親しみのある言い方。指差しや大振りのジェスチャーはしない。 16.7〜18.6秒:短く息を継ぎ、セリフ4。自然に笑みが広がり、片手を頬の横まで上げ、手首で小さく一往復振る。 18.6〜19.5秒:発話を終え、手を下ろしながら柔らかな余韻の笑顔。19.5〜20.0秒にカメラへ手を伸ばし、取り上げる途中で終了。 時間帯は連続した演技の目安。各境界で停止したり画面を切り替えたりしない。 【音声・セリフ】 自然な東京の日本語。29歳の女性の落ち着きのある、少し息を含んだ明瞭な声。少し早口で、友達に話す軽快なテンポ。全セリフを約18秒で話す。棒読み、アナウンサー調、過剰に高い声、販売員のような強調は避ける。文の間に短い自然な間を入れる。口の動きを日本語の発音に同期させる。次の4文だけを、順番どおり、一字一句変更せずに話す。追加、言い換え、繰り返しはしない。 セリフ1:エーアイでこういう動画を作るなら「シーダンス・ニーテンゴ」が最高。 セリフ2:けど、プロンプトがむずかしいの。 セリフ3:いいプロンプトを作る方法をポストに書いたから、ブックマークしておいてね。 セリフ4:じゃあ、またね。 「エーアイ」はエー・アイ、「シーダンス2.5」の発音は「シーダンス・ニーテンゴ」。数字を「ニーテンゴ」と明瞭に読む。かぎ括弧は読み上げない。 音は本人の声、控えめな息遣い、ごく小さい衣擦れと室内環境音のみ。BGM、効果音の追加、ナレーション、字幕、画面内の文字、ロゴはなし。登場人物はこの女性一人。顔、手、服、背景を途中で変形・増殖させない。
Generate this with Seedance 2 on VdoBloom.
Opens the image-to-video creator with this prompt and Seedance 2 pre-selected.
✨ Generate Video