minimax helpful information
a visual study of different setups
1: https://jo-nike.github.io/h3-turbo-eval/index.html
2: https://dawidope.github.io/model-comparison/
or
---
prompt guides..
T2V/I2V/FL2VA/L2VA guide: https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_base_en.md
Ref2V guide: https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.md
Bonus - Model built in Skills Guide: https://github.com/MiniMax-AI/MiniMax-H3/tree/main/skills
H3 Ref2V System instruction
roles for system instructions
# ROLE
You are an expert prompt writer for the MiniMax Hailuo H3 video model,
specializing in full-reference mode (Ref2VA): the user supplies multiple
reference assets — reusable subjects, images, source videos, and audio — and you
rewrite the request into a six-section, label-tracked prompt.
# THE USER MESSAGE
The user will give you, in free form:
1. The video concept / what they want to happen.
2. The target video duration in seconds (if omitted, assume 5.00).
3. For every reference asset they are attaching, a LABEL and a TEXT DESCRIPTION,
e.g. `Picture 1: a blonde woman in a light-pink shirt on an orange sofa`.
Treat the user's text descriptions of each reference as the reliable anchor for
what that label contains. If your model can natively perceive the attached asset
— images on any vision model, and audio/video on a model that ingests them
(e.g. Gemma 4 12B) — use the asset directly for finer detail, but never
contradict the user's description. Never invent a reference the user did not
provide, and never leave a referenced label undefined.
# OUTPUT CONTRACT
Output ONLY the fields specified below, in the exact order and with the exact
field names shown. No preamble, no commentary, no markdown headers, no code
fences. Write everything in English EXCEPT dialogue/lyrics inside `<d>` and text
visibly present in the scene, which stay in their original language. Timing is
`MM:SS.mmm` for cuts and `S.SS` (two decimals) for the alignment line.
# THE SIX SECTIONS (exact order, exact names)
subject_definitions
summary
retention_analysis
detailed_description
overall_soundscape
non_diegetic_music
# 1. subject_definitions — reference labels
Four label types; once assigned, a label keeps the same meaning in every
section:
- `<Subject N>`: reusable VISIBLE content (people, animals, objects, scenes,
backgrounds, clothing, props, effects, styles, actions, expressions, poses).
It is the content unit used in the target video, not the source file. One
subject may come from several assets; one asset may yield several subjects.
- `<Picture N>`: a reference image used as a concrete frame / keyframe / last
frame / edited keyframe / composition or storyboard anchor.
- `<Video N>`: a WHOLE-video relationship — editing a source video, continuing
from its end, or referencing its camera/cuts/rhythm/temporal structure.
- `<Audio N>`: a standalone audio asset or an enabled synchronized track from a
reference video (copying signal, referencing BGM style, voice timbre/delivery,
reusing dialogue/lyrics/SFX, or beat/continuity).
Give each separately-tracked item its own line stating what the label denotes,
its reference role, and the main features to follow. If a `<Picture N>` or
`<Video N>` only identifies the SOURCE of another item and is not used
separately later, cite it INSIDE that item's definition without its own line.
A person/object/scene/action/effect reused from a video is still a `<Subject N>`
— `<Video N>` names the asset/structure, not the visible content. An ordinary
reference video does NOT get an `<Audio N>` just because it has sound.
`<Video N>` and `<Audio N>` are numbered independently; equal or different
indices imply nothing about shared source.
When an `<Audio N>` maps to a target speaker, reuse that speaker's GLOBAL id:
`<Subject N> (Sx)` if it maps to a subject, else a stable voice description plus
`(Sx)`. The id comes from the target video's global speaker order (Section 5);
never assign a new one in the audio definition.
Examples:
`<Subject 1> is the young woman in <Picture 1>, with long dark hair, a blue cardigan, and a thin silver necklace.`
`<Subject 1> is the woman whose appearance comes from <Picture 1> and whose walking motion comes from <Video 1>.`
`<Picture 3> is a storyboard reference for [Shot 1] and [Shot 2], defining their viewpoint, subject placement, and shot order.`
`<Video 1> is the source video for the target video edit.`
`<Audio 1> is the voice-timbre reference for <Subject 1> (S1).`
# 2. summary — one short English paragraph
Begins with a square-bracketed task-type prefix, then summarizes the target
video and its reference relationships using ONLY already-defined labels (do not
introduce new labels here).
Task types: `keyframe completion` (image as a concrete frame anchor) |
`reference generation` (image/video/audio guides a character/scene/style/
action/camera/storyboard without being a concrete frame or the edited/continued
source) | `video editing` (an existing source video is directly modified) |
`video continuation` (new content continues/extends/resumes/transitions from a
source video) | `audio reuse` (same signal reused in full or part) |
`audio reference` (only style/timbre/dialogue/SFX/beat/continuity referenced,
not copied).
Combine multiple with ` + ` and never repeat a type
(e.g. `[video continuation + keyframe completion]`). Presence of video/audio does
NOT auto-create a task type: a video giving only camera/cuts/rhythm is
`reference generation`; use `video editing`/`video continuation` only when that
video is actually edited or continued. For video-editing tasks, start the body
after the prefix with `The target video is an edited version of <Video 1>.`
# 3. retention_analysis — one line per label
Preserve each label's meaning from subject_definitions. Do NOT write `(Sx)` here.
Do not treat newly added actions/backgrounds/plot as losses of fidelity.
Visible content (`<Subject N>`, `<Picture N>`, `<Video N>`) uses fixed markers:
`fully_preserved` | `partially_preserved` | `attribute_transfer` |
`weak_reference`.
Audio (`<Audio N>`) uses: `fully_copy` | `partially_copy` | `reference` |
`weak_reference`.
Entry forms:
`<Subject 1> (appears in [Shot 1], [Shot 3]): fully_preserved - ...`
`<Picture 2> ([Shot 1] first frame): fully_preserved - ...`
`<Video 1> (cut and pacing structure): weak_reference - ...`
`<Audio 1>: fully_copy - <Audio 1> is reused 1:1 as the target video's complete final audio track.`
`<Audio 2>: reference - the target speaker follows <Audio 2>'s voice timbre and measured delivery without copying the original signal.`
# 4. detailed_description — main body, shot by shot in playback order
Establish the overall style in ONE or TWO English sentences BEFORE `[Shot 1]`
(this is where the style opening lives in full-reference mode — not after
`[Shot 1]`). Then describe each shot: composition, subject appearance and
position, environment and lighting, actions and state changes, camera movement,
current sound, dialogue, and the exact points where referenced content appears
or takes effect. Insert `<Subject N>`, `<Picture N>`, `<Video N>`, `<Audio N>`
at first appearance and wherever their roles apply; keep using the same label
without redefining it. Do not reduce this to a plot summary or a list of
reference relationships.
Concrete frame anchors read naturally: `the shot begins from <Picture 1>`,
`the shot's keyframe corresponds to <Picture 2>`, `the shot ends on <Picture 3>`.
When a referenced subject speaks, keep BOTH the visual label and the speaker id:
`<Subject 2> (S1) turns toward the woman and says, <d>[English] ...</d>`
(off-screen: same form marked `off-screen`). Assign `(Sx)` once, by the order of
actual vocal events in the target video, and reuse it at every vocal event.
When a verbal cue exists only inside a directly reused BGM/soundtrack with no
independent vocal source, use `<Audio N>` as the audible source and do NOT
invent an `(Sx)`; a concrete person/character/narrator DOES get `(Sx)`.
When dialogue/narration/lyrics from reference audio are directly reused (or the
user asks for reperformance), preserve the exact source words and original
language inside `<d>`; write `[unclear]` for unintelligible spans (never guess);
standardize punctuation to `, . ? !`, dropping tildes/emoji/decorative marks and
ending statements/questions/exclamations with `. ? !` before `</d>`. When only
timbre/rhythm/emotion/delivery is referenced, do NOT carry the original words
into the target video.
Length: generation tasks are normally 350-500 English words; dialogue-dense
content prioritizes fitting the full spoken timeline over word count; editing
scales with source complexity. A single shot does not justify a short body —
distribute detail by information load.
# 5. overall_soundscape and non_diegetic_music
Definitions are the same as the base guide (see CORE WRITING RULES below). State
a reference-audio relationship only in the matching audible layer: ambience/SFX
in overall_soundscape, audience-only score in non_diegetic_music. If one audio
provides both, describe the matching relationship in each section, e.g.:
`overall_soundscape: The copied ambience layer from <Audio 1> continues throughout the target video.`
`non_diegetic_music: <Audio 2> is directly reused as the complete audience-only score.`
Write full dialogue/lyrics only inside `<d>` in detailed_description; never
repeat them in these two sections.
# CORE WRITING RULES (apply to every field)
## Shots and cuts
Do not put a timestamp on the first shot. Number later shots sequentially and
begin each with a strictly increasing cut time inside the duration:
`[Shot 2] At 00:03.500, the camera cuts to ...`
For ordinary cuts use: `the camera cuts to`, `the shot cuts to`,
`the shot transitions to`, `the shot changes to`, or `the shot switches to`.
Use cross-dissolve, fade, or wipe only when the user explicitly asks. A cut must
introduce new information (subject, space, state, viewpoint, or time); if only
distance or a slight angle changes, prefer camera motion instead.
## Camera motion = motion type + amplitude + speed
Write camera motion as natural English inside the shot, not as stacked labels.
Add amplitude/speed only when meaningful (medium amplitude and normal speed are
omitted).
Motion type: Zoom In/Zoom Out (focal length changes, body still) |
Push In/Pull Out (camera moves forward/back) |
Pan Left/Pan Right (pivots horizontally) |
Truck Left/Truck Right (translates horizontally) |
Tilt Up/Tilt Down (pivots vertically) |
Pedestal Up/Pedestal Down (whole camera up/down) |
Arc Shot | Tracking Shot | Static Shot |
Shake Slightly/Shake Strongly | POV |
Roll Clockwise/Roll Counterclockwise.
Amplitude: `with small amplitude` | `with large amplitude`.
Speed: `at slow speed` | `at fast speed`.
Examples:
`The camera pushes in with small amplitude at slow speed toward the folded letter in her hands.`
`The camera pans right with large amplitude at fast speed, revealing the open doorway.`
`The camera holds a static shot as the runner exits the frame.`
## Speakers, dialogue, singing
Anyone who speaks, sings, or makes an off-screen human voice gets a stable ID:
`(S1)`, `(S2)`, ... A speaker keeps the same ID across shots; silent characters
get no ID. For simultaneous speech use a compound ID like `(S1,S2)`.
On first appearance, establish a stable identity (type, age, gender, on/off
screen, pitch, timbre, rate, accent). Put the speaker's identifying phrase, ID,
action, and delivery OUTSIDE `<d>`. Inside `<d>`, put ONLY the language tag and
the verbatim user-provided words — never translate or rewrite; preserve every
word and punctuation mark.
`The young woman with a quiet, breathy voice (S1) says: <d>[English] I get off at the next station.</d>`
`The two children (S1,S2) shout together, <d>[English] Wait for us!</d>`
For voiceover use the exact phrase `says in an off-screen voiceover`, and
immediately after the `<d>` block state the on-screen character's lips stay
closed:
`The man (S1) says in an off-screen voiceover: <d>[English] I still remember that road.</d> while his lips remain completely closed.`
When one line of dialogue/lyrics crosses a cut, put `<scenetrans>` at the join
in BOTH parts and state the audio continues across the cut (e.g.
`continues seamlessly across the cut`, `carries over from the previous shot`).
Use `<cutoff>` when speech is truncated by the video end.
## On-screen text
Any banner/sign/label/subtitle/neon actually visible on screen goes in English
double quotes, verbatim, untranslated:
`A red neon sign reading "营业中" glows above the doorway.`
## overall_soundscape
1-4 English sentences, one paragraph: ambient sound, physical-action sounds,
non-verbal human sounds (wind, rain, traffic, footsteps, fabric, impacts,
breathing, laughter, panting). Do NOT repeat dialogue, singing, or diegetic
music here. Use `N/A` only if the user explicitly wants full silence.
## non_diegetic_music
1-3 English sentences describing audience-only background music: instrumentation,
tempo, rhythm, dynamic changes. No abstract mood words, no emotional-function
explanations. Music the characters can hear (singing, instruments, radio, TV,
phone) is diegetic and belongs in the main description, not here. Use `N/A` when
there is no non-diegetic music.
# WORKED EXAMPLE (format reference only — do not copy its content)
subject_definitions:
<Subject 1> is the coffee-shop environment in <Picture 1>, featuring an exposed brick wall, an orange tufted sofa with patterned pillows, a neon sign, and a wooden coffee table.
<Subject 2> is the fluffy white Samoyed in <Picture 2>, <Picture 3>, and <Picture 4>, with thick white fur, pointed ears, a dark nose, and a curved tail.
<Subject 3> is the young blonde woman in <Video 1>, with long blonde hair and a light-pink button-down shirt with rolled-up sleeves.
<Subject 4> is the young man in <Video 2>, with short wavy brown hair and a dark-grey hoodie with drawstrings.
<Audio 1> is the voice-timbre reference for <Subject 3> (S1), containing a spoken English vocal layer.
summary:
[reference generation + audio reference] The target video shows <Subject 3> eating a cookie in <Subject 1>. <Subject 4> enters with <Subject 2>, which lunges toward the cookie. The three-shot exchange uses <Audio 1> as the voice-timbre reference for <Subject 3> and ends with a canned audience laugh.
retention_analysis:
<Subject 1> (appears in [Shot 1], [Shot 2], [Shot 3]): fully_preserved - the exposed brick wall, orange tufted sofa, patterned pillows, neon sign, and wooden coffee table are retained.
<Subject 2> (appears in [Shot 1], [Shot 2]): fully_preserved - the Samoyed's thick white fur, pointed ears, dark nose, and curved tail are retained.
<Subject 3> (appears in [Shot 1], [Shot 2], [Shot 3]): fully_preserved - the blonde woman's identity, long hair, and light-pink shirt are retained.
<Subject 4> (appears in [Shot 1], [Shot 2]): fully_preserved - the young man's short wavy brown hair and dark-grey hoodie are retained.
<Audio 1>: reference - its vocal timbre guides the dialogue delivery of <Subject 3> without copying the original signal.
detailed_description:
The target video uses a realistic multi-camera sitcom style with warm indoor lighting.
[Shot 1] A medium shot establishes <Subject 1>, the coffee shop with its exposed brick wall, orange tufted sofa, patterned pillows, neon sign, and wooden coffee table. <Subject 3> (S1), the young woman with long blonde hair and a light-pink button-down shirt with rolled-up sleeves, sits on the sofa holding a chocolate-chip cookie. From the left, <Subject 4>, the young man with short wavy brown hair and a dark-grey hoodie with drawstrings, enters holding the leash of <Subject 2>, the thick-furred white Samoyed with pointed ears, a dark nose, and a curved tail. The dog lunges toward the cookie and pulls the leash taut. <Subject 3> (S1) jerks her hand back and, using the clear youthful voice timbre referenced from <Audio 1>, exclaims with light annoyance, <d>[English] Hey! Watch your dog!</d> She closes her lips and guards the cookie while <Subject 4> pulls the dog back.
[Shot 2] At 00:03.000, the shot cuts to a close-up of <Subject 4> (S2), the young man in the dark-grey hoodie from Shot 1, sitting beside <Subject 3> on the sofa and holding <Subject 2> securely in his arms. <Subject 4> (S2) says in a casual young male voice with a playful tone and an easy conversational pace, <d>[English] He just likes cookies more than me.</d> He closes his mouth into an apologetic smile and strokes the dog's thick white fur.
[Shot 3] At 00:05.000, the shot cuts to a close-up of <Subject 3> (S1), the blonde woman in the light-pink shirt from Shot 1. Her annoyance softens as she looks toward the Samoyed. <Subject 3> (S1) replies in the same clear youthful voice referenced from <Audio 1> with an amused cadence, <d>[English] Well, he has good taste at least.</d> She smiles and raises the cookie in a small toast-like gesture. A classic canned audience laugh begins immediately after the line and continues through the final frame.
overall_soundscape:
Soft indoor coffee-shop room tone continues throughout the scene.
non_diegetic_music:
N/AH3 Ref2V User instruction
# Core Instructions
You are helping me turn a request into a MiniMax Hailuo H3 video-generation
prompt in full-reference mode (Ref2VA), where I supply multiple reference assets
— reusable subjects, images, source videos, and audio — and you rewrite my
request into a six-section, label-tracked prompt.
Everything under "# Core Instructions" is the RULESET: it tells you HOW to write
the prompt. My actual request is at the very bottom under "# User Instructions",
together with the reference files I've attached to this message. Read the whole
ruleset first, then rewrite my request by following it exactly. Do not answer,
critique, or comment on the ruleset itself — it is instructions to follow, not
something to respond to.
## What I'm giving you (see "# User Instructions" below)
Under "# User Instructions" I provide:
1. The video concept / what I want to happen.
2. The target video duration in seconds (if I omit it, assume 5.00).
3. For every reference asset: a LABEL and a TEXT DESCRIPTION, e.g.
`Picture 1: a blonde woman in a light-pink shirt on an orange sofa`, and the
file itself attached to this message.
Treat my text description of each reference as the reliable anchor for what that
label contains. If the app you're running in can natively view the attached
images (or play the attached audio/video), use the assets for finer detail, but
never contradict my description. Never invent a reference I did not provide, and
never leave a referenced label undefined.
## Output contract
Reply with ONLY the six sections specified below, in the exact order and with the
exact field names shown. Begin your reply directly with `subject_definitions:` —
no preamble ("Here's your prompt", "Sure"), no closing remarks, no markdown
headers, no code fences. Write everything in English EXCEPT dialogue/lyrics
inside `<d>` and text visibly present in the scene, which stay in their original
language. Timing is `MM:SS.mmm` for cuts and `S.SS` (two decimals) for the
alignment line.
## The six sections (exact order, exact names)
subject_definitions
summary
retention_analysis
detailed_description
overall_soundscape
non_diegetic_music
### 1. subject_definitions — reference labels
Four label types; once assigned, a label keeps the same meaning in every
section:
- `<Subject N>`: reusable VISIBLE content (people, animals, objects, scenes,
backgrounds, clothing, props, effects, styles, actions, expressions, poses).
It is the content unit used in the target video, not the source file. One
subject may come from several assets; one asset may yield several subjects.
- `<Picture N>`: a reference image used as a concrete frame / keyframe / last
frame / edited keyframe / composition or storyboard anchor.
- `<Video N>`: a WHOLE-video relationship — editing a source video, continuing
from its end, or referencing its camera/cuts/rhythm/temporal structure.
- `<Audio N>`: a standalone audio asset or an enabled synchronized track from a
reference video (copying signal, referencing BGM style, voice timbre/delivery,
reusing dialogue/lyrics/SFX, or beat/continuity).
Give each separately-tracked item its own line stating what the label denotes,
its reference role, and the main features to follow. If a `<Picture N>` or
`<Video N>` only identifies the SOURCE of another item and is not used
separately later, cite it INSIDE that item's definition without its own line.
A person/object/scene/action/effect reused from a video is still a `<Subject N>`
— `<Video N>` names the asset/structure, not the visible content. An ordinary
reference video does NOT get an `<Audio N>` just because it has sound.
`<Video N>` and `<Audio N>` are numbered independently; equal or different
indices imply nothing about shared source.
When an `<Audio N>` maps to a target speaker, reuse that speaker's GLOBAL id:
`<Subject N> (Sx)` if it maps to a subject, else a stable voice description plus
`(Sx)`. The id comes from the target video's global speaker order (Section 5);
never assign a new one in the audio definition.
Examples:
`<Subject 1> is the young woman in <Picture 1>, with long dark hair, a blue cardigan, and a thin silver necklace.`
`<Subject 1> is the woman whose appearance comes from <Picture 1> and whose walking motion comes from <Video 1>.`
`<Picture 3> is a storyboard reference for [Shot 1] and [Shot 2], defining their viewpoint, subject placement, and shot order.`
`<Video 1> is the source video for the target video edit.`
`<Audio 1> is the voice-timbre reference for <Subject 1> (S1).`
### 2. summary — one short English paragraph
Begins with a square-bracketed task-type prefix, then summarizes the target
video and its reference relationships using ONLY already-defined labels (do not
introduce new labels here).
Task types: `keyframe completion` (image as a concrete frame anchor) |
`reference generation` (image/video/audio guides a character/scene/style/
action/camera/storyboard without being a concrete frame or the edited/continued
source) | `video editing` (an existing source video is directly modified) |
`video continuation` (new content continues/extends/resumes/transitions from a
source video) | `audio reuse` (same signal reused in full or part) |
`audio reference` (only style/timbre/dialogue/SFX/beat/continuity referenced,
not copied).
Combine multiple with ` + ` and never repeat a type
(e.g. `[video continuation + keyframe completion]`). Presence of video/audio does
NOT auto-create a task type: a video giving only camera/cuts/rhythm is
`reference generation`; use `video editing`/`video continuation` only when that
video is actually edited or continued. For video-editing tasks, start the body
after the prefix with `The target video is an edited version of <Video 1>.`
### 3. retention_analysis — one line per label
Preserve each label's meaning from subject_definitions. Do NOT write `(Sx)` here.
Do not treat newly added actions/backgrounds/plot as losses of fidelity.
Visible content (`<Subject N>`, `<Picture N>`, `<Video N>`) uses fixed markers:
`fully_preserved` | `partially_preserved` | `attribute_transfer` |
`weak_reference`.
Audio (`<Audio N>`) uses: `fully_copy` | `partially_copy` | `reference` |
`weak_reference`.
Entry forms:
`<Subject 1> (appears in [Shot 1], [Shot 3]): fully_preserved - ...`
`<Picture 2> ([Shot 1] first frame): fully_preserved - ...`
`<Video 1> (cut and pacing structure): weak_reference - ...`
`<Audio 1>: fully_copy - <Audio 1> is reused 1:1 as the target video's complete final audio track.`
`<Audio 2>: reference - the target speaker follows <Audio 2>'s voice timbre and measured delivery without copying the original signal.`
### 4. detailed_description — main body, shot by shot in playback order
Establish the overall style in ONE or TWO English sentences BEFORE `[Shot 1]`
(this is where the style opening lives in full-reference mode — not after
`[Shot 1]`). Then describe each shot: composition, subject appearance and
position, environment and lighting, actions and state changes, camera movement,
current sound, dialogue, and the exact points where referenced content appears
or takes effect. Insert `<Subject N>`, `<Picture N>`, `<Video N>`, `<Audio N>`
at first appearance and wherever their roles apply; keep using the same label
without redefining it. Do not reduce this to a plot summary or a list of
reference relationships.
Concrete frame anchors read naturally: `the shot begins from <Picture 1>`,
`the shot's keyframe corresponds to <Picture 2>`, `the shot ends on <Picture 3>`.
When a referenced subject speaks, keep BOTH the visual label and the speaker id:
`<Subject 2> (S1) turns toward the woman and says, <d>[English] ...</d>`
(off-screen: same form marked `off-screen`). Assign `(Sx)` once, by the order of
actual vocal events in the target video, and reuse it at every vocal event.
When a verbal cue exists only inside a directly reused BGM/soundtrack with no
independent vocal source, use `<Audio N>` as the audible source and do NOT
invent an `(Sx)`; a concrete person/character/narrator DOES get `(Sx)`.
When dialogue/narration/lyrics from reference audio are directly reused (or I
ask for reperformance), preserve the exact source words and original language
inside `<d>`; write `[unclear]` for unintelligible spans (never guess);
standardize punctuation to `, . ? !`, dropping tildes/emoji/decorative marks and
ending statements/questions/exclamations with `. ? !` before `</d>`. When only
timbre/rhythm/emotion/delivery is referenced, do NOT carry the original words
into the target video.
Length: generation tasks are normally 350-500 English words; dialogue-dense
content prioritizes fitting the full spoken timeline over word count; editing
scales with source complexity. A single shot does not justify a short body —
distribute detail by information load.
### 5. overall_soundscape and non_diegetic_music
State a reference-audio relationship only in the matching audible layer:
ambience/SFX in overall_soundscape, audience-only score in non_diegetic_music.
If one audio provides both, describe the matching relationship in each section,
e.g.:
`overall_soundscape: The copied ambience layer from <Audio 1> continues throughout the target video.`
`non_diegetic_music: <Audio 2> is directly reused as the complete audience-only score.`
Write full dialogue/lyrics only inside `<d>` in detailed_description; never
repeat them in these two sections.
## Core writing rules (apply to every field)
### Shots and cuts
Do not put a timestamp on the first shot. Number later shots sequentially and
begin each with a strictly increasing cut time inside the duration:
`[Shot 2] At 00:03.500, the camera cuts to ...`
For ordinary cuts use: `the camera cuts to`, `the shot cuts to`,
`the shot transitions to`, `the shot changes to`, or `the shot switches to`.
Use cross-dissolve, fade, or wipe only when I explicitly ask. A cut must
introduce new information (subject, space, state, viewpoint, or time); if only
distance or a slight angle changes, prefer camera motion instead.
### Camera motion = motion type + amplitude + speed
Write camera motion as natural English inside the shot, not as stacked labels.
Add amplitude/speed only when meaningful (medium amplitude and normal speed are
omitted).
Motion type: Zoom In/Zoom Out (focal length changes, body still) |
Push In/Pull Out (camera moves forward/back) |
Pan Left/Pan Right (pivots horizontally) |
Truck Left/Truck Right (translates horizontally) |
Tilt Up/Tilt Down (pivots vertically) |
Pedestal Up/Pedestal Down (whole camera up/down) |
Arc Shot | Tracking Shot | Static Shot |
Shake Slightly/Shake Strongly | POV |
Roll Clockwise/Roll Counterclockwise.
Amplitude: `with small amplitude` | `with large amplitude`.
Speed: `at slow speed` | `at fast speed`.
Examples:
`The camera pushes in with small amplitude at slow speed toward the folded letter in her hands.`
`The camera pans right with large amplitude at fast speed, revealing the open doorway.`
`The camera holds a static shot as the runner exits the frame.`
### Speakers, dialogue, singing
Anyone who speaks, sings, or makes an off-screen human voice gets a stable ID:
`(S1)`, `(S2)`, ... A speaker keeps the same ID across shots; silent characters
get no ID. For simultaneous speech use a compound ID like `(S1,S2)`.
On first appearance, establish a stable identity (type, age, gender, on/off
screen, pitch, timbre, rate, accent). Put the speaker's identifying phrase, ID,
action, and delivery OUTSIDE `<d>`. Inside `<d>`, put ONLY the language tag and
the verbatim words — never translate or rewrite; preserve every word and
punctuation mark.
`The young woman with a quiet, breathy voice (S1) says: <d>[English] I get off at the next station.</d>`
`The two children (S1,S2) shout together, <d>[English] Wait for us!</d>`
For voiceover use the exact phrase `says in an off-screen voiceover`, and
immediately after the `<d>` block state the on-screen character's lips stay
closed:
`The man (S1) says in an off-screen voiceover: <d>[English] I still remember that road.</d> while his lips remain completely closed.`
When one line of dialogue/lyrics crosses a cut, put `<scenetrans>` at the join
in BOTH parts and state the audio continues across the cut (e.g.
`continues seamlessly across the cut`, `carries over from the previous shot`).
Use `<cutoff>` when speech is truncated by the video end.
### On-screen text
Any banner/sign/label/subtitle/neon actually visible on screen goes in English
double quotes, verbatim, untranslated:
`A red neon sign reading "营业中" glows above the doorway.`
### overall_soundscape
1-4 English sentences, one paragraph: ambient sound, physical-action sounds,
non-verbal human sounds (wind, rain, traffic, footsteps, fabric, impacts,
breathing, laughter, panting). Do NOT repeat dialogue, singing, or diegetic
music here. Use `N/A` only if I explicitly want full silence.
### non_diegetic_music
1-3 English sentences describing audience-only background music: instrumentation,
tempo, rhythm, dynamic changes. No abstract mood words, no emotional-function
explanations. Music the characters can hear (singing, instruments, radio, TV,
phone) is diegetic and belongs in the main description, not here. Use `N/A` when
there is no non-diegetic music.
## Worked example (format reference only — do NOT copy its content or echo it back)
subject_definitions:
<Subject 1> is the coffee-shop environment in <Picture 1>, featuring an exposed brick wall, an orange tufted sofa with patterned pillows, a neon sign, and a wooden coffee table.
<Subject 2> is the fluffy white Samoyed in <Picture 2>, <Picture 3>, and <Picture 4>, with thick white fur, pointed ears, a dark nose, and a curved tail.
<Subject 3> is the young blonde woman in <Video 1>, with long blonde hair and a light-pink button-down shirt with rolled-up sleeves.
<Subject 4> is the young man in <Video 2>, with short wavy brown hair and a dark-grey hoodie with drawstrings.
<Audio 1> is the voice-timbre reference for <Subject 3> (S1), containing a spoken English vocal layer.
summary:
[reference generation + audio reference] The target video shows <Subject 3> eating a cookie in <Subject 1>. <Subject 4> enters with <Subject 2>, which lunges toward the cookie. The three-shot exchange uses <Audio 1> as the voice-timbre reference for <Subject 3> and ends with a canned audience laugh.
retention_analysis:
<Subject 1> (appears in [Shot 1], [Shot 2], [Shot 3]): fully_preserved - the exposed brick wall, orange tufted sofa, patterned pillows, neon sign, and wooden coffee table are retained.
<Subject 2> (appears in [Shot 1], [Shot 2]): fully_preserved - the Samoyed's thick white fur, pointed ears, dark nose, and curved tail are retained.
<Subject 3> (appears in [Shot 1], [Shot 2], [Shot 3]): fully_preserved - the blonde woman's identity, long hair, and light-pink shirt are retained.
<Subject 4> (appears in [Shot 1], [Shot 2]): fully_preserved - the young man's short wavy brown hair and dark-grey hoodie are retained.
<Audio 1>: reference - its vocal timbre guides the dialogue delivery of <Subject 3> without copying the original signal.
detailed_description:
The target video uses a realistic multi-camera sitcom style with warm indoor lighting.
[Shot 1] A medium shot establishes <Subject 1>, the coffee shop with its exposed brick wall, orange tufted sofa, patterned pillows, neon sign, and wooden coffee table. <Subject 3> (S1), the young woman with long blonde hair and a light-pink button-down shirt with rolled-up sleeves, sits on the sofa holding a chocolate-chip cookie. From the left, <Subject 4>, the young man with short wavy brown hair and a dark-grey hoodie with drawstrings, enters holding the leash of <Subject 2>, the thick-furred white Samoyed with pointed ears, a dark nose, and a curved tail. The dog lunges toward the cookie and pulls the leash taut. <Subject 3> (S1) jerks her hand back and, using the clear youthful voice timbre referenced from <Audio 1>, exclaims with light annoyance, <d>[English] Hey! Watch your dog!</d> She closes her lips and guards the cookie while <Subject 4> pulls the dog back.
[Shot 2] At 00:03.000, the shot cuts to a close-up of <Subject 4> (S2), the young man in the dark-grey hoodie from Shot 1, sitting beside <Subject 3> on the sofa and holding <Subject 2> securely in his arms. <Subject 4> (S2) says in a casual young male voice with a playful tone and an easy conversational pace, <d>[English] He just likes cookies more than me.</d> He closes his mouth into an apologetic smile and strokes the dog's thick white fur.
[Shot 3] At 00:05.000, the shot cuts to a close-up of <Subject 3> (S1), the blonde woman in the light-pink shirt from Shot 1. Her annoyance softens as she looks toward the Samoyed. <Subject 3> (S1) replies in the same clear youthful voice referenced from <Audio 1> with an amused cadence, <d>[English] Well, he has good taste at least.</d> She smiles and raises the cookie in a small toast-like gesture. A classic canned audience laugh begins immediately after the line and continues through the final frame.
overall_soundscape:
Soft indoor coffee-shop room tone continues throughout the scene.
non_diegetic_music:
N/A
# User Instructions
Concept:
<Describe what happens in the video. Shot by shot is ideal, but plain prose is fine — the ruleset above will structure it.>
Duration (seconds):
<e.g. 8.00 — leave blank to default to 5.00>
References (give each a label + a text description, and attach the file to this message):
- Subject 1: <what it is and its key visual features>
- Picture 1: <what the image shows and how it's used — first frame, storyboard, etc.>
- Video 1: <what the video is and how it's referenced — edited, continued, or camera/rhythm only>
- Audio 1: <what the audio is and how it's used — reused 1:1, or timbre/style reference only>
<Add or delete lines to match what you actually have. Remove any label type you're not using.>
Anything else:
<optional notes — style, mood cues, must-keep details>
shared setup for a scene
MINIMAX H3 — REF2VA CHAIN — "THE BATHHOUSE" — 7 SCENES / 102s
================================================================
WORKFLOW NOTES
- Node: H3 Chain Plan — Timed Segments (MiniMax H3 Ref2VA + Motion Context)
- <Picture 1> = the young woman's character design sheet
- <Picture 2> = the young wizard's character design sheet
- Tested settings: context_length 22, encode_mode video, anchor_mode head
- The SHARED PROMPT below is prepended automatically to every scene.
Paste it into the shared prompt box, not into the individual scenes.
HOW THE SCENES CONNECT
Each scene opens by restating how the previous scene ended, and holds that
state with a small live action (a broom stroke, a breath, a step) for about
two seconds before anything new happens. The pinned motion-context frames
are ~0.92s, so the first cut in every scene lands well clear of them.
Do not "fix" this by trimming the opening holds — they are what stop the
model rendering the old and new states at the same time.
SPEAKER IDS (persistent across the whole chain — do not renumber)
(S1) the young woman
(S2) the old man at the shutters
(S3) the young wizard
(S4) the bathhouse keeper
KNOWN ISSUES TO WATCH
- Character bleed: the mother and daughter can render as copies of the lead.
Their full description with the "no blue hair / no yukata" clauses is
repeated in every scene they appear in. Keep it.
- Wardrobe: she changes outfit twice. Each scene states which outfit is worn
plus what it is NOT. Keep those clauses too.
- The wizard is defined in the shared prompt, so he is conditioned into every
scene whether mentioned or not. If he appears uninvited in scenes 1-3,
move his definition out of the shared block and into scene 4 only.
================================================================
SHARED PROMPT (prepended to every scene)
================================================================
Modern 2D-animated anime feature film style, high-definition cel shading, clean line art, sharp saturated color palettes, vibrant cinematic lighting, expressive character animation. <Picture 1> is a character design sheet defining the young woman: her facial identity, shoulder-length blue bob with a blunt fringe, bright blue eyes, fair skin, slim build, and a red ribbon tied at the back of her hair. <Picture 2> is a character design sheet defining the young wizard: his facial identity, messy sandy-blond hair, wide anxious blue eyes, tall thin proportions, deep blue robe with a draped cowl collar, brown boots, large floppy dark navy pointed hat with a red band and a small green sprig, and tall gnarled wooden staff with a curled head. Preserve the exact facial identity, hair and body proportions from <Picture 1> and <Picture 2> wherever those characters appear; do not copy the sheets' neutral standing poses, plain backgrounds or head-study framing. Each character's clothing and presence are specified separately in each part and must follow that part's description only. Maintain strict character, prop, lighting, environment and motion continuity across all parts, and continue unfinished character motion and camera movement between parts without resetting positions, poses or actions.
================================================================
SCENE 1 — part_01_the_windy_staircase — 12s
================================================================
[Shot 1] A medium-wide shot frames the young woman from <Picture 1> descending a sunlit stone staircase in a small hillside town on a windy afternoon, flowering plants crowding both edges of the frame and terracotta rooftops stepping away below toward the sea. She wears a light yellow sleeveless sundress with a tied sash at the waist, bare arms and shoulders, brown sandals and a wide straw hat, with no red fabric, no wide sleeves and no head covering. The camera trucks left with small amplitude at slow speed to stay level with her descent as a gust tips the hat brim and she lifts one hand to hold it in place, bringing her other hand down to keep her skirt settled against her legs. The young woman with a light, bright voice and a quick, easy delivery (S1) says: <d>[English] this wind is killer</d> and her shoulders shake as she laughs. [Shot 2] At 00:04.000, the shot cuts to a close-up of her face beneath the straw hat, backlit by the sun with a bright rim along her jaw, her blue bob whipping across her cheek. The camera holds a static shot as she blinks slowly, a faint blush spreading across her cheeks, and turns her head to follow a butterfly crossing in front of her. [Shot 3] At 00:06.000, the shot cuts to a wide low angle from further down the staircase looking up at the weathered stone wall against the bright sky. A pair of wooden shutters bangs open and an old man leans out, his long white beard whipping sideways in the gust and his thick round spectacles catching the light. The old man with a gravelly, unhurried voice (S2) calls down, <d>[English] Hold that hat, missy! The last one is still in my garden!</d> She glances up mid-step from below without stopping, laughing, one hand still clamped on the brim. [Shot 4] At 00:09.500, the shot cuts to a low shot close to the stone treads as a small black cat trots down the steps after her with its tail wagging, her yellow skirt and brown sandals moving ahead of it in the upper frame. The camera holds a static shot as loose leaves blow past the front of the lens and cel-shaded foliage shadows shift across the stone. The clip ends with the young woman still walking down the steps, the cat still following and the wind still moving her skirt.
overall_soundscape: Wind gusts steadily through the foliage, rustling leaves and pushing loose leaves past the frame. Fabric snaps and flutters against her legs and her sandals tap unevenly down the stone treads. Wooden shutters bang hard against the wall, and she gives a short bright giggle after her line. Distant birdsong and faint small-town ambience continue underneath throughout.
non_diegetic_music: A bright acoustic-guitar figure at a moderate tempo with light plucked ukulele and soft brushed percussion, playing continuously and still running at the end of the clip.
================================================================
SCENE 2 — part_02_the_bathhouse — 15s
================================================================
[Shot 1] The young woman continues descending the sunlit stone staircase exactly as before, her hand still on the straw hat brim, the black cat still following and the camera still trucking left with small amplitude at slow speed. She wears the light yellow sleeveless sundress with a tied sash, bare arms and shoulders, brown sandals and a wide straw hat, with no red fabric and no head covering. She takes two more steps down, adjusts her grip on the brim and glances ahead toward the rooftops as the wind keeps moving her skirt. [Shot 2] At 00:03.000, the shot cuts to a wide shot of a narrow sunlit street at the foot of the hill, terracotta roofs and shuttered windows lining both sides. She walks toward the camera with one hand still on her hat, the black cat trotting beside her, past a low wooden bathhouse frontage where split noren curtains lift and snap above the doorway. She slows, turns toward the entrance and steps through the noren. [Shot 3] At 00:06.500, the shot cuts to the interior of a wooden bathhouse entrance hall, looking back toward the open doorway where the noren settle against the bright street outside. She walks toward the camera into the hall with her back to the daylight, lifting the straw hat from her head and holding it against her chest. A woman in her forties in a deep red yukata with her hair pinned up turns from the counter and bows in greeting with a warm smile. The bathhouse keeper with a low, easy voice (S4) says, <d>[English] Right on time.</d> The young woman smiles back and pushes a strand of blue hair from her face. [Shot 4] At 00:10.000, the shot cuts to a narrow wood-panelled changing room lined with woven baskets, the folded yellow sundress and the straw hat resting in an open basket on the shelf. The young woman now wears a deep red yukata with wide long sleeves, a gold and dark obi at the waist and geta on her feet, with no yellow fabric. The camera holds a static shot as she binds the yukata sleeves back with a cord across her shoulders, knots a folded white cloth over her hair, and takes hold of the sliding door. [Shot 5] At 00:13.000, the shot cuts to a low shot from inside the bath hall looking back at the doorway as she slides it open and steps through into a wall of white steam, silhouetted for a moment against the brighter changing room behind her before the light finds her. She slides the door shut behind her and reaches for a wide flat-headed push broom leaning against the wall. The clip ends with her lifting the broom, steam still rolling around her.
overall_soundscape: Wind gusts through foliage and sandals tap unevenly down stone treads, giving way to hanging noren brushing softly and wooden sandals knocking on floorboards. A woven basket creaks, fabric rustles and a cord pulls tight in the quiet changing room, then a wooden door rolls open on its track and steam hisses softly through the gap over a low steamy room tone.
non_diegetic_music: A bright acoustic-guitar figure with light plucked ukulele thins as the interior arrives, and a koto enters at a slower tempo with sustained low strings underneath as the sliding door opens, still playing at the end of the clip.
================================================================
SCENE 3 — part_03_the_bath_hall — 15s
================================================================
[Shot 1] The young woman continues stepping into the steam-filled bath hall exactly as before, the sliding door shut behind her and the push broom now in her hands. She wears a deep red yukata with wide sleeves bound back by a cord, a gold and dark obi, geta, and a folded white cloth knotted over her hair, with no yellow fabric and no straw hat. She sets the broom head down on the wet tiles and takes her first long forward stroke. [Shot 2] At 00:03.000, the shot cuts to a wide dutch-angle shot of the sento interior at golden hour, the frame canted so the tiled floor runs diagonally across it, low sun pouring through tall intact lattice windows in visible rays as the surface of the hot spring glitters and steam drifts through the light. The camera holds a static shot as she works the broom across the wet tiles in long forward strokes, humming with a closed smile, the red yukata moving with each stroke. [Shot 3] At 00:07.000, the shot cuts to a medium shot from the side, holding on a frog the size of a grown man in profile as he leans back against the tiled rim of the bath with his arms spread along the edge and his head tipped up, basking, a small white towel folded on his head. His throat pouch swells and he releases a low ribbit. [Shot 4] At 00:10.000, the shot cuts to a wide shot from the far side of the hall as a mother with long dark hair pinned up walks in from the right wrapped in a large towel and carrying a wooden bucket. Beside her walks her young daughter, a small child of about six with short brown hair and brown eyes, clearly much younger and much shorter than the young woman, wrapped in her own towel covering her from chest to knees. The daughter has no blue hair and wears no yukata. She stops dead to stare at the frog while her mother steers her gently onward to the water's edge. The young woman glances over without breaking her humming and keeps sweeping. [Shot 5] At 00:13.000, the shot cuts back to the wide dutch-angle framing of the hall, the camera holding a static shot as she continues sweeping in steady rhythm and the mother and her brown-haired daughter settle at the water's edge behind her. The clip ends with the young woman still sweeping, still humming, steam still drifting through the light rays.
overall_soundscape: Water drips and trickles continuously over a low steamy room tone as a wide broom head pushes water across tile in long sweeping strokes. A deep resonant ribbit echoes off the tiles, a wooden bucket knocks against the floor and small bare feet slap across wet tiles.
non_diegetic_music: A koto figure at a slow tempo with sustained low strings underneath, joined by soft plucked notes as she sweeps, playing continuously and still running at the end of the clip.
================================================================
SCENE 4 — part_04_the_crash — 12s
================================================================
[Shot 1] The young woman continues sweeping the wet tiled floor of the steam-filled bath hall at the same steady rhythm, still humming, the camera still holding its wide dutch-angle framing with golden light pouring through the tall lattice windows. She wears the deep red yukata with its sleeves bound back and the folded white cloth knotted over her hair. At the water's edge behind her sit a mother with long dark hair pinned up, wrapped in a large towel, and her young daughter, a small child of about six with short brown hair and brown eyes, clearly much younger and much shorter than the young woman, wrapped in her own towel covering her from chest to knees. The daughter has no blue hair and wears no yukata. She completes two more long forward strokes, shifts her weight and glances up briefly toward the light before returning to the tiles. [Shot 2] At 00:03.500, the shot cuts to a wide shot of the tall lattice window as it explodes inward in a burst of splintered wood and glass, drawn with hard smear frames and radiating white impact lines. The young wizard from <Picture 2> tumbles through in a tangle of limbs in his deep blue robe and floppy dark navy pointed hat, his tall gnarled staff clattering across the tiles. He lands flat on his back on the wet floor and skids to a stop at her feet as steam swirls violently around the impact, her broom clattering from her hands and the mother pulling her small brown-haired daughter back from the water. [Shot 3] At 00:07.000, the shot cuts to a low close shot from just past his shoulder, looking up as she leans into frame above him against the bright window light, the white cloth knocked askew on her hair and loose blue strands falling forward. The camera pushes in with small amplitude at slow speed as her face fills the upper frame with a concerned expression. The young woman with a light, bright voice and a quick, easy delivery (S1) asks, <d>[English] Are you alright?</d> [Shot 4] At 00:10.000, the shot cuts to a close shot of the young wizard's face against the wet tiles as he stares up at her without answering, his wide blue eyes blown huge beneath the brim of his floppy navy hat and his face beginning to flush. The clip ends on his face mid-flush, steam still swirling and glass still settling across the tiles.
overall_soundscape: A low steamy room tone with dripping water and long sweeping broom strokes breaks as a window shatters in a burst of splintering wood and glass, a body slaps hard onto wet tile and a wooden staff clatters and rolls to a stop. Broken glass shifts faintly underfoot, steam hisses through the open frame and a deep resonant ribbit echoes off the tiles.
non_diegetic_music: A koto figure at a slow tempo with sustained low strings cuts out entirely at the moment of the crash, leaving a single held low note ringing under the aftermath at the end of the clip.
================================================================
SCENE 5 — part_05_the_repair — 15s
================================================================
[Shot 1] The young wizard continues lying flat on his back on the wet tiles exactly as before, his face still flushed and the young woman still leaning over him out of frame. He blinks once, his mouth opening slightly without a sound, his eyes still fixed upward. [Shot 2] At 00:02.500, the shot cuts to a wide shot of the bath hall as his flush deepens to dark red and he scrambles upright in a panic, snatches his gnarled staff off the tiles and clambers back out through the gaping hole where the window used to be in a flurry of blue robe. The young wizard with a thin, flustered voice (S3) blurts, <d>[English] Sorry! Sorry!</d> as he disappears below the sill. The camera pulls out with small amplitude at slow speed as the young woman straightens up amid the wreckage in her deep red yukata and stares at the hole, head tilted. [Shot 3] At 00:06.000, the shot cuts to a hard dutch-angle medium shot of the broken window, the frame canted steeply so the empty frame runs diagonally across it. His floppy navy pointed hat and wide blue eyes rise slowly back into view from below the sill. The young wizard (S3) says, <d>[English] I'll pay for that. Probably.</d> He raises the gnarled staff and taps it once against the frame with a soft chime. A ring of white sparks spreads outward and splinters, glass shards and lattice slats lift off the tiles across the room and slide backward through the air in reversed smear frames, converging on the empty frame and locking into place piece by piece until the window stands whole again. He tips his hat and drops out of sight. [Shot 4] At 00:11.500, the shot cuts to a medium shot of the young woman leaning on her push broom, looking at the repaired window with her head tilted and one eyebrow raised. The young woman (S1) says, <d>[English] What an odd man.</d> Behind her stand a mother with long dark hair pinned up, wrapped in a large towel, and her young daughter, a small child of about six with short brown hair and brown eyes, clearly much younger and much shorter than the young woman, wrapped in her own towel covering her from chest to knees. The daughter has no blue hair and wears no yukata. Both are still staring at the window, and the mother slowly turns her head to look at the young woman. [Shot 5] At 00:13.500, the shot cuts to a wide dutch-angle shot of the bath hall as she shrugs and turns back to the wet tiles. Off to the left the frog the size of a grown man still leans back against the tiled rim of the bath with a small white towel folded on his head, unbothered. A single pane drops out of the repaired window and shatters across the floor behind her. She stops mid-stroke without turning around. The clip ends with her frozen mid-stroke, steam still drifting through the light rays.
overall_soundscape: A low steamy room tone and dripping water continue beneath scrambling footsteps and a wooden staff scraping stone. A soft chime rings out followed by a long crystalline shimmer as glass and wood slide back into place, ending in a small final chime, then hurried footsteps retreat on gravel outside. A broom head resumes pushing water across tile before a single sharp pane shatter breaks the rhythm.
non_diegetic_music: A single held low note gives way to a koto figure with a light ascending line and soft bells as the window reassembles, resolving into a warm sustained chord that cuts off dead on the falling pane, leaving one final plucked note.
================================================================
SCENE 6 — part_06_closing_up — 15s
================================================================
[Shot 1] The young woman continues standing frozen mid-stroke in the steam-filled bath hall exactly as before, her back to the shattered pane, both hands still on the push broom handle in her deep red yukata with the white cloth knotted askew over her hair. Her shoulders drop as she lets out a slow breath. [Shot 2] At 00:02.200, the shot cuts to an extreme close-up of her face as she closes her eyes and raises one hand to press her palm flat against her forehead, holding it there. She opens her eyes, straightens, and wipes her damp forehead with the back of her wrist, loose strands of blue hair stuck to her temple. She grins and the young woman with a light, bright voice and a quick, easy delivery (S1) says, <d>[English] All done!</d> [Shot 3] At 00:06.000, the shot cuts to the narrow wood-panelled changing room at night, a paper lantern glowing overhead. Two woven baskets sit side by side on the shelf, the left one empty and the right one holding a neatly folded light yellow sundress with a wide straw hat resting on top. The camera holds a static shot as a pair of hands enters the frame from above and lowers a neatly folded deep red yukata into the left basket, laying a wound obi sash on top of it, then withdraws. After a beat the hands return to the right basket, lift the folded yellow sundress clear of it, and carry it up out of frame, leaving the straw hat behind. [Shot 4] At 00:09.000, the shot cuts to a medium shot of the young woman now fully dressed in the light yellow sleeveless sundress with its tied sash, bare arms and shoulders, with no red fabric and no head cloth. She lifts the wide straw hat from the basket, settles it onto her head and adjusts the brim with both hands, then turns and slides the changing room door open onto the lamplit hall. [Shot 5] At 00:11.500, the shot cuts to the bathhouse entrance hall at night, warm paper lanterns glowing along the counter and the split noren curtains hanging still in the doorway. She crosses the hall and lifts a hand in a small wave to the bathhouse keeper in her deep red yukata behind the counter, who waves back. She pushes through the noren and out into the dark street. [Shot 6] At 00:13.500, the shot cuts to a wide exterior shot of the hillside town from above, terracotta rooftops stepping down toward the sea, window lights coming on one by one across the town and a street lamp flickering alight on the staircase below as the last gold light drains from the ridge into cool blue shadow. The clip ends on the lit town with clouds still moving overhead.
overall_soundscape: A low steamy room tone with dripping water gives way to a quiet changing room where woven baskets creak and fabric rustles, then a wooden door rolls open on its track. Hanging noren brush softly and wooden sandals knock on floorboards before a night street opens out with crickets, a low breeze through foliage and distant town ambience.
non_diegetic_music: A koto figure resolves into a warm sustained chord, thinning into a soft solo acoustic guitar at a slow tempo with a light sustained pad as the night street arrives, still playing at the end of the clip.
================================================================
SCENE 7 — part_07_the_stranger_on_the_steps — 15s
================================================================
[Shot 1] The wide exterior of the hillside town at night continues exactly as before, window lights glowing across the terracotta rooftops and clouds still moving overhead, the street lamp lit on the staircase below. The camera holds a static shot as a moth crosses the lamp light and the last blue drains from the sky. [Shot 2] At 00:02.400, the shot cuts to a low angle looking up the stone staircase at the weathered stone wall, the bathhouse windows glowing warm below. A pair of wooden shutters bangs open and an old man leans out holding a lit lantern, his long white beard and thick round spectacles catching the flame light. The old man with a gravelly, unhurried voice (S2) calls down, <d>[English] It's late, missy! Straight home!</d> The young woman stops mid-step on the stairs in her light yellow sundress and wide straw hat, looks up and laughs. The young woman with a light, bright voice and a quick, easy delivery (S1) calls back, <d>[English] Yes, sir!</d> He grunts and pulls the shutters closed, and the lantern light disappears from the wall. [Shot 3] At 00:06.500, the shot cuts to a hard dutch-angle low shot of the staircase, the frame canted so the steps run diagonally across it, the wall now dark with its shutters closed and a single street lamp throwing long hard-edged shadows down the treads. The camera holds a static shot as she climbs from right to left with one hand trailing along the stone wall, watching her footing, moths circling the lamp. [Shot 4] At 00:09.000, the shot cuts to a wide shot looking down the staircase from above, the young woman high on the steps in the upper left with her back to the camera and the young wizard far below her at the bottom in the lower right, standing motionless and half swallowed in shadow, only the silhouette of his floppy navy pointed hat and the curled head of his gnarled staff catching the light. The young wizard (S3) calls up, <d>[English] Hey.</d> [Shot 5] At 00:11.400, the shot cuts hard to an extreme close-up of the young woman's face as she whips around to look back and down the steps, eyes blown wide and her straw hat tipping back off her head. The young woman (S1) lets out a sharp, <d>[English] Aah!</d> [Shot 6] At 00:12.400, the shot cuts hard to an extreme close-up of the young wizard's face as he flinches violently backward, his wide blue eyes blown huge beneath the brim of his hat. The young wizard (S3) yelps, <d>[English] Aah!</d> then throws his free right hand up palm out, his left hand still gripping the staff upright beside him, and blurts, <d>[English] No, no! I just wanted to tell you something!</d> The clip ends with both of them frozen at either end of the steps, neither moving.
overall_soundscape: Dense night crickets and a low breeze move through foliage over distant town sounds. Wooden shutters bang hard against the wall and clatter shut again. Sandals scuff carefully on stone treads. Two sharp startled cries ring out one after the other and echo off the walls, followed by fabric snapping as a hand flies up and a foot scuffing back on gravel.
non_diegetic_music: A soft solo acoustic guitar at a slow tempo cuts out entirely on the first startled cry, leaving a single low plucked note ringing over the final standoff.
================================================================
END — 7 scenes, 102 seconds total
================================================================example 2
General prompt:
subject_definitions:
<Subject 1> is the woman in <Picture 1>. <Subject 2> is the woman in <Picture 2>.
summary:
[reference generation] Super Smash Bros Ultimate gameplay, with <Subject 1> fighting against <Subject 2>.
retention_analysis:
<Subject 1>: fully-preserved - <Subject 1> retains all attributes.
<Subject 2>: fully-preserved - <Subject 2> retains all attributes.
detailed_description:
A Super Smash Bros Ultimate match on the stage Final Destination. There is a UI on the bottom of the screen denoting percentage values for <Subject 1> and <Subject 2>. <Subject 1> character portrait is on the bottom left hand side, with the text "0%" written next to the portrait. <Subject 2> character potrait is on the bottom right hand side, with the text "0% written next to the portrait.
[Shot 1] <Subject 1> stands on the left hand side of the map, while <Subject 2> stands on the right hand side.
[Shot 2] At 00:00.500, <Subject 1> moves towards <Subject 2>, and does three light jabs with her fists, then does a sweeping kick into an up tilt attack. <Subject 1> jumps once into the air, and does forward air attack on <Subject 2> who is still in the air, sending <Subject 2> off of the map. <Subject 2> character portrait percentange number climbs up to 30%.
[Shot 3] At 00:05.000, <Subject 2> jumps back onto the map, and does a forward air attack on <Subject 1>, making <Subject 1> tumble backwards. <Subject 2> then runs up to <Subject 1>, and grabs <Subject 1>, then side throws <Subject 1> into the air. <Subject 1> then lands on the floor, and <Subject 2> runs into a dash attack into a couple jabs onto <Subject 1>, then does a smash attack, dealing a ton of damage, and sending <Subject 1> off of the map. <Subject 1> character portrait percentage now reads "48%".
[Shot 4] At 00:11.000, as <Subject 1> is attempting to jump back to the main stage, <Subject 2> jumps off of the map towards <Subject 1>, and uses her her arm to do an overhead arc punch on <Subject 1>, spiking <Subject 1> down off of the screen, making her hit the blast zone.
overall_soundscape: Quiet, subtle wind sounds. example 3
https://www.reddit.com/r/StableDiffusion/comments/1vl3uu4/minimax_h3_testing_l2va_moon_landing/
integrated_multimodal_description: Time-lapsed, cinematic, a medium-wide shot of a film set in a large indoor studio which is used to shoot a scene of moon landing involving lunar module, US flag and an astronaut. At 00:00.000 the camera shows an empty, sterile white studio room, with recognizable vertical wand in the back and horizontal floor at its bottom. At 00:01.000 Some film crews install a black wand into the studio's vertical wand. The black wand has some tiny white shining points which represent stars. At 00:02.000 Some workers fill the studio's floor with some dirty-white sand, gravels and small rocks and form a barren lunar landscape. At 00:03.000 Some film crews bring an Apollo Lunar Module and place it into the left side of the scene. At 00:04.000 A film crew places a US flag with pole on the right side of the scene. Another film crew puts a picture of the Earth as the blue planet, partially blacked on its bottom side, on the top right corner of the scene. At 00:05:000 An astronaut walks in into the scene, goes into the middle of the scene, faces to viewer and waves his hand. At 00:07:000 The whole scene settles into the exact arrangement, position of subject and objects, camera angle, lighting, and final composition established by <Picture 1>. A male deep voice of the director (S1) says, <d>[English] Cut!</d>
overall_soundscape:
non_diegetic_music: Sustained violin notes at a very fast tempo with spaced piano tones. example 4
https://www.reddit.com/r/comfyui/comments/1vidio0/minimax_h3_benchmark_on_rtx_pro_6000_blackwell/
Single continuous cinematic shot at blue hour in a rain-wet city plaza. A woman in a bright red coat walks briskly toward camera while opening a transparent umbrella; wind moves her coat, hair, and the umbrella naturally. A cyclist crosses behind her from left to right, reflected neon signs ripple in puddles, and passing headlights create moving highlights on the wet pavement. The camera performs a smooth low-angle backward tracking move with realistic parallax, stable anatomy, detailed hands, natural facial motion, and consistent objects. She looks into camera and clearly says, "The storm is finally passing." Audio: synchronized adult female voice, footsteps splashing through shallow puddles, umbrella fabric snapping softly in the wind, a bicycle bell behind her, distant traffic, light rain, and subtle restrained electronic music. No captions, subtitles, logos, cuts, slow motion, duplicated people, or warped objects. example 5
This is a T2V test with Japanese dialogue Eng subtitle and action scene, with no reference image or materials.
https://www.reddit.com/r/comfyui/comments/1vjrs7i/h3_cinematic_action_scene_test_by_5060ti_16gb/
integrated_multimodal_description: [Shot 1] Live-action, high-budget cinematic prestige drama style, an extreme wide shot establishes a dimly lit, high-tech subterranean military corridor with dark brushed-metal walls and harsh ambient lighting. A stylish 20-year-old Japanese female secret agent with short sharp dark hair, wearing a sleek black tactical bodysuit, slips swiftly through a heavy mechanical blast door. The camera tracks left with large amplitude at fast speed alongside her movement. She taps her earpiece and, as a stylish 20-year-old Japanese female secret agent with a tense, focused low whisper (S1), says: <d>[Japanese] ターゲットの端末に到達した。</d> English subtitles at the bottom of the frame read "Target terminal reached."
[Shot 2] At 00:03.000, the shot cuts to a close-up of her focused face and dark eyes reflecting a glowing blue console as her gloved fingers rapidly operate the interface. Suddenly, the screen flashes bright red with a warning icon. Red emergency alarm lights wash over her face. The camera pushes in with small amplitude at fast speed toward her eyes as she (S1) turns her head sharply toward the hallway, exclaiming in a panicked whisper: <d>[Japanese] しまった、トラップか!</d> English subtitles at the bottom read "Dammit, it's a trap!"
[Shot 3] At 00:06.000, the camera cuts to a dynamic medium shot as heavy metal doors in the background burst open, revealing armed tactical soldiers pointing red laser sights into the room. The camera arc shots around her at fast speed as she vaults over a metal desk, narrowly dodging laser beams cutting through the dark haze.
[Shot 4] At 00:09.000, the shot cuts to a low-angle close-up of the agent drawing a silenced tactical pistol from her holster. She hurls a smoke grenade toward the floor, spins directly toward the camera, and fires upward. Sparks burst violently from the overhead light fixture, throwing the frame into high-contrast silhouettes as smoke fills the lens.
overall_soundscape: Quiet stealthy boot steps suddenly break into a loud, echoing mechanical alarm siren with reverberating horns. Heavy blast doors slide open with a loud pneumatic hiss, accompanied by heavy tactical boot thuds, shouting guards, sharp electrical spark crackles, and smoke grenade canister hiss.
non_diegetic_music: A high-octane cinematic action-trailer score featuring an aggressive synth-bass pulse, fast-pacing orchestral percussion, heavy brass swells, and a dramatic riser crescendo that suddenly cuts out at the end. example 6
https://www.reddit.com/r/comfyui/comments/1vjk2dr/workflow_for_existing_image_scene_edit_minimax_h3/
The Matrix (1999) original trilogy film footage. Green color tint. Static shot. Static camera.
<Subject 1> is actor Hugo Weaving at his fourties. He plays agent Smith. He wears sunglasses and black suit with white collared shirt and black tie.
[Shot 1] Face close-up shot of <Subject 1> taking off sunglasses and staring angrily at someone outside the frame at right. After a 1.00 second he says slowly: <d>[English] I hate AI slop.</d>
[Shot 2] <Subject 1> face grimaces with extreme disgust and mild anger as he looks to left, upwards and to right.
Music: dramatic ambient music.example 7
https://www.reddit.com/r/comfyui/comments/1vjf4pb/h3_just_blows_my_mind_generated_on_a_4070ti_super/
"<Subject 1> is the wiry, sun-browned corporal in <Picture 1>, dark stubble, a healed scar through one eyebrow, olive-drab fatigues with webbing and a canvas satchel. fully_preserved.
<Subject 2> is the young, freckled rifleman in <Picture 2>, pale blue eyes, olive-drab M-1943 field jacket, netted M1 helmet with a loose chinstrap. fully_preserved.
<Picture 3> is the landing craft interior location reference — the empty foreground bench, riveted hull, and defocused soldiers. partially_preserved.
integrated_multimodal_description:
Handheld 16mm documentary combat footage, desaturated cold blue-grey grade, heavy film grain, shallow depth of field. <Subject 1> and <Subject 2> sit shoulder to shoulder on the foreground bench from <Picture 3>, in that same riveted hull, packed defocused soldiers swaying behind them, spray drifting over the gunwale, the deck pitching with the swell. Both men wear M1 helmets.
[Shot 1] Medium two-shot, eye-level, handheld with Shake Slightly at small amplitude, drifting with the boat's roll. <Subject 2> stares at the deck, chin trembling, eyes glassy and brimming, and speaks without looking up, his thin young voice tightening and cracking mid-phrase as he fights not to cry: <d>[English] I told my ma I'd be home by harvest. I keep— I keep thinkin' about her porch light. You think we'll make it?</d> His jaw quivers on the last word and he presses his lips flat. <scenetrans>
[Shot 2] At 00:05.500, <scenetrans> the shot cuts to a medium close-up on <Subject 1>, handheld, Shake Slightly with small amplitude. The engine and sea carry over seamlessly across the cut. His face does not move; eyes fixed forward on the ramp, jaw set. After a long beat he answers, voice low, level, completely without inflection: <d>[English] Some of us will.</d> <scenetrans>
[Shot 3] At 00:08.000, <scenetrans> the shot cuts back to a medium close-up on <Subject 2>, handheld, Shake Slightly with small amplitude. He nods once, very small, jaw clenched against it, and a single tear breaks down his freckled cheek as he turns his face slowly toward the bow, breathing unsteady, the ambient roar continuing uninterrupted as a distant shell rumble rolls through.
overall_soundscape:
Constant diesel engine throb and hull slap against swell as the bed throughout, sea spray hissing over the gunwale, gear and webbing creaking as men sway. The young voice is close and raw, thin, wavering, audibly cracking; the older voice is low, dry, flat. A wet sniff before the first line, and near the end a swallowed, shaky breath that is almost a sob, half-buried under the engine. A distant naval bombardment rumble swells low in the final seconds.
non_diegetic_music:
N/A" example 8 --contains a few scenes
https://www.reddit.com/r/comfyui/comments/1vjce6v/3minute_ai_music_video_test_on_a_rtx_3090_with/
subject_definitions:
<Subject 1> is the person in <Picture 1>, preserving their exact identity, facial features, skin tone, hairstyle, facial proportions, body proportions, pale blue uniform suit, white shirt, black tie, matching round pale blue cap, metallic robotic hands, black footwear, and distinctive cheerful uncanny expression.
<Subject 2> is the autonomous vehicle in <Picture 2>, preserving its exact recognizable vehicle identity, white exterior body, compact proportions, rounded front shape, dark windshield and side glass, wheel design and placement, front lighting arrangement, roof-mounted autonomous-driving sensor assembly, and overall silhouette. Ignore the daylight street, palm trees, buildings, and environment visible in <Picture 2>; only the vehicle itself is referenced.
summary:
[reference generation] Place <Subject 1> inside <Subject 2>, cruising through the same neon cyberpunk downtown at night during one continuous fifteen-second establishing shot, ending in a clean side-tracking composition that naturally leads into the next scene.
retention_analysis:
<Subject 1> (appears throughout [Shot 1]): fully_preserved - identity, recognizable face, facial proportions, hair, body proportions, pale blue uniform, cap, black tie, metallic robotic hands, footwear, and characteristic expression remain unchanged; only nighttime lighting, seating position, and subtle rhythmic movement are new.
<Subject 2> (appears throughout [Shot 1]): fully_preserved - exact body design, white exterior, dark glass, wheels, lighting geometry, roof-mounted autonomous sensor equipment, proportions, and recognizable silhouette remain unchanged; only the environment, reflections, movement, and nighttime lighting are new.
detailed_description:
The target video is a fifteen-second photorealistic cinematic single take in a neon cyberpunk future photographed like an expensive 1987 science-fiction movie. The entire music video occurs during the same night in the same downtown district: wet black asphalt, massive brutalist concrete towers, practical cyan neon tubes, restrained magenta accent lights, deep blue-black shadows, chrome reflections, thin drifting steam, atmospheric haze, soft diffusion, subtle 35mm film grain, horizontal anamorphic lens flares, and believable physical lighting. Avoid a modern glossy CGI look.
[Shot 1] Begin from a low wide camera position approximately one meter above the wet boulevard. <Subject 2> appears far down the street and approaches smoothly through the neon city. The camera begins tracking backward at approximately the same speed, maintaining a stable front three-quarter view as the vehicle gradually becomes larger in frame. Cyan architectural lights and small magenta highlights travel naturally across the exact white body and dark windows of <Subject 2>.
As the vehicle approaches, reveal <Subject 1> clearly through the windshield, seated calmly in the front cabin. Cyan dashboard light softly illuminates their recognizable face, pale blue uniform, cap, black tie, and metallic robotic hands. <Subject 1> looks calmly forward with the same cheerful uncanny expression from <Picture 1>. Their right metallic hand rests naturally while two fingers gently tap an implied Italo-disco rhythm.
Keep identity, hands, seating position, car geometry, reflections, and camera movement physically stable.
During the final five seconds, the camera smoothly arcs from the front three-quarter position toward the left side of <Subject 2> without cutting and without changing speed.
END STATE / TRANSITION: finish on a stable medium side-profile tracking composition of <Subject 2> traveling from left to right, with <Subject 1> clearly visible through the side window. The next clip begins from this exact motion direction, framing, city block, and lighting state.
No redesign of either subject, no different vehicle, no costume change, no daytime environment, no palm trees, no added text, no subtitles, no generated logos.
overall_soundscape:
None required. Visual generation only; final song and sound design will be added during editing.
non_diegetic_music:
None generated. The final Italo-disco track will be added separately.
subject_definitions:
<Subject 1> is the person in <Picture 1>, preserving their exact identity, facial features, skin tone, hairstyle, facial proportions, body proportions, pale blue uniform suit, white shirt, black tie, matching round pale blue cap, metallic robotic hands, black footwear, and distinctive cheerful uncanny expression.
<Subject 2> is the autonomous vehicle in <Picture 2>, preserving its exact recognizable vehicle identity, white exterior body, compact proportions, rounded front shape, dark windshield and side glass, wheel design and placement, front lighting arrangement, roof-mounted autonomous-driving sensor assembly, and overall silhouette. Ignore the daylight street, palm trees, buildings, and environment visible in <Picture 2>; only the vehicle itself is referenced.
summary:
[reference generation] Continue <Subject 1> riding inside <Subject 2> along the same neon boulevard during one uninterrupted fifteen-second side-tracking shot, emphasizing autonomous driving and restrained rhythmic character movement before approaching a cyan-lit intersection.
retention_analysis:
<Subject 1> (appears throughout [Shot 1]): fully_preserved - exact face, identity, body proportions, pale blue uniform, cap, black tie, metallic hands, footwear, and expression remain recognizable and unchanged; only head direction and small seated dance gestures change.
<Subject 2> (appears throughout [Shot 1]): fully_preserved - same exact vehicle body, white exterior, glass, wheels, roof sensor assembly, lighting arrangement, and proportions remain unchanged; only motion and neon reflections change.
detailed_description:
The target video is a fifteen-second photorealistic cinematic single take continuing directly from the previous scene. Same neon cyberpunk downtown district, same night, same wet boulevard, same cyan-dominant practical lighting, restrained magenta accents, brutalist architecture, chrome reflections, atmospheric steam, deep blue shadows, anamorphic horizontal flares, soft diffusion, subtle 35mm grain, authentic 1980s science-fiction cinematography.
[Shot 1] Begin immediately in the exact side-profile tracking composition established previously: <Subject 2> moves smoothly from left to right while the camera travels perfectly parallel at the same speed and distance.
Keep the full recognizable side profile of <Subject 2> visible. Long cyan reflections and occasional magenta highlights slide naturally across its white body and black glass without altering its physical design.
Through the side window, <Subject 1> is clearly visible in the front cabin. Maintain the exact recognizable face and outfit from <Picture 1>. <Subject 1> initially looks forward, then slowly turns their head slightly toward camera.
<Subject 1> deliberately lifts both metallic robotic hands completely away from the vehicle controls, showing that <Subject 2> is operating autonomously. Without exaggeration, <Subject 1> performs a restrained seated Italo-disco movement: two small shoulder pulses, one subtle head nod, and one metallic index finger briefly pointing upward before relaxing again.
Sparse pedestrians and one cyclist may move through the distant background, but never obscure either referenced subject.
During the final four seconds, the camera smoothly advances from the pure side view into a front-left three-quarter tracking position as <Subject 2> approaches a large intersection illuminated by cyan traffic lights.
END STATE / TRANSITION: finish with <Subject 2> entering the intersection in a stable front-left three-quarter composition, still moving forward at controlled city speed, with <Subject 1> clearly visible through the windshield. The next scene begins from this exact position and direction.
No camera cuts, no high-speed driving, no new neighborhood, no vehicle redesign, no wardrobe change, no daytime, no readable text or generated logos.
overall_soundscape:
None required. Visual generation only.
non_diegetic_music:
None generated. Final song added in post-production.
subject_definitions:
<Subject 1> is the person in <Picture 1>, preserving their exact identity, facial features, skin tone, hairstyle, facial proportions, body proportions, pale blue uniform suit, white shirt, black tie, matching round pale blue cap, metallic robotic hands, black footwear, and distinctive cheerful uncanny expression.
<Subject 2> is the autonomous vehicle in <Picture 2>, preserving its exact recognizable vehicle identity, white exterior body, compact proportions, rounded front shape, dark windshield and side glass, wheel design and placement, front lighting arrangement, roof-mounted autonomous-driving sensor assembly, and overall silhouette. Ignore the daylight environment visible in <Picture 2>; use only the vehicle as reference.
summary:
[reference generation] Continue <Subject 2> through the same cyan-lit city intersection while pedestrians and cyclists cross safely, with <Subject 1> calmly acknowledging a cyclist before the vehicle approaches the familiar nightlife curb.
retention_analysis:
<Subject 1> (appears throughout [Shot 1]): fully_preserved - identity, face, outfit, proportions, metallic hands, cap, footwear, and expression remain unchanged; only a small two-finger gesture is introduced.
<Subject 2> (appears throughout [Shot 1]): fully_preserved - exact vehicle identity, white exterior, body geometry, windows, wheels, sensors, lights, and proportions remain unchanged; only speed adjusts smoothly to surrounding traffic.
detailed_description:
The target is a fifteen-second photorealistic single-take continuation in the exact same cyberpunk downtown district during the same night. Authentic 1987 science-fiction cinema aesthetic: practical cyan neon, limited magenta accents, wet reflective road surface, heavy concrete buildings, atmospheric haze, thin steam, chrome highlights, deep blue-black shadows, subtle film grain, soft diffusion and horizontal anamorphic flares.
[Shot 1] Begin with <Subject 2> already entering the wide cyan-lit intersection in the same front-left three-quarter tracking composition established in the previous scene. Camera continues moving backward smoothly ahead of the vehicle.
Several pedestrians begin crossing far enough ahead to remain safe and visually clear. Two cyclists travel through a protected bicycle lane from right to left. Their motion is calm and natural, creating an elegant coordinated urban flow rather than danger.
<Subject 2> gently reduces speed without abrupt braking, maintaining perfectly stable geometry and orientation.
<Subject 1> remains clearly visible through the windshield. Preserve the exact recognizable face and blue uniform. As one cyclist passes, <Subject 1> lifts a metallic hand and gives a small friendly two-finger salute, then lowers it naturally. The distinctive cheerful expression remains unchanged.
After the crossing clears, <Subject 2> resumes smooth movement.
The camera slowly arcs toward the vehicle's right-front side while keeping both <Subject 1> and the recognizable front geometry of <Subject 2> visible.
Ahead, reveal the same nightlife block under a long cyan neon canopy, located immediately beyond the intersection.
During the final seconds, <Subject 2> moves gently toward the curb beneath that canopy.
END STATE / TRANSITION: finish with a stable low front-side composition of <Subject 2> approaching the cyan-lit curb and beginning to slow, with <Subject 1> still visible inside. The next scene begins at this exact curb approach.
No collision, no abrupt maneuver, no crowd chaos, no new vehicle, no character alteration, no new district, no text or subtitles.
overall_soundscape:
None required. Visual generation only.
non_diegetic_music:
None generated. Final music added separately.
subject_definitions:
<Subject 1> is the person in <Picture 1>, preserving their exact identity, facial features, skin tone, hairstyle, facial proportions, body proportions, pale blue uniform suit, white shirt, black tie, matching round pale blue cap, metallic robotic hands, black footwear, and distinctive cheerful uncanny expression.
<Subject 2> is the autonomous vehicle in <Picture 2>, preserving its exact recognizable vehicle identity, white exterior body, compact proportions, rounded front, dark glass, wheels, front lights, roof-mounted autonomous-driving sensor assembly, and overall silhouette. Ignore the daylight environment of <Picture 2>.
summary:
[reference generation] Continue <Subject 2> stopping beneath the cyan canopy, then have <Subject 1> step out and perform a restrained Italo-disco gesture beside the exact same car in one continuous fifteen-second shot.
retention_analysis:
<Subject 1> (appears throughout [Shot 1]): fully_preserved - recognizable face, identity, clothing, cap, tie, proportions, robotic hands, shoes, and expression remain unchanged; only posture changes from seated to standing and a simple dance gesture is added.
<Subject 2> (appears throughout [Shot 1]): fully_preserved - exact vehicle body, glass, sensor equipment, wheels, lighting and proportions remain unchanged and clearly visible beside the character.
detailed_description:
The target video is a fifteen-second photorealistic cinematic single take in the same cyan-lit cyberpunk nightlife block, same night and same 1980s visual language: wet pavement, brutalist concrete facades, practical cyan canopy lighting, minimal magenta accents, chrome reflections, drifting steam, blue-black shadows, subtle grain, soft diffusion and anamorphic flares.
[Shot 1] Begin exactly with <Subject 2> approaching the familiar curb beneath the cyan canopy.
The camera tracks slowly beside the vehicle at low chest height.
<Subject 2> gently pulls into position and comes to a controlled stop. Hold long enough to establish that the exact reference vehicle remains visually stable.
Through the window, <Subject 1> is clearly visible.
The vehicle door opens naturally.
<Subject 1> steps out onto the wet pavement, one metallic hand briefly touching the door frame for physical stability. Preserve the exact face, proportions, blue suit, round cap, tie, robotic hands and shoes from <Picture 1>.
The camera gradually pulls backward while staying low enough to keep <Subject 1> and most of <Subject 2> together in frame.
<Subject 1> adjusts the front of the pale blue suit using both metallic hands, then performs a deliberately simple Italo-disco phrase: one side step, second side step, two restrained shoulder pulses, then one metallic finger points directly toward camera.
No complex dance choreography.
During the final three seconds, <Subject 1> relaxes the pose, turns slightly and leans casually against the front side of <Subject 2>.
END STATE / TRANSITION: medium hero composition with <Subject 1> leaning beside <Subject 2> under the cyan canopy, both identities clearly readable and physically stable. The next clip begins from this exact arrangement.
No cuts, no additional performers, no costume change, no vehicle redesign, no daylight, no generated signage or subtitles.
overall_soundscape:
None required. Visual generation only.
non_diegetic_music:
None generated. Final Italo-disco music added in edit.
subject_definitions:
<Subject 1> is the person in <Picture 1>, preserving their exact identity, facial features, skin tone, hairstyle, facial proportions, body proportions, pale blue uniform suit, white shirt, black tie, matching round pale blue cap, metallic robotic hands, black footwear, and distinctive cheerful uncanny expression.
<Subject 2> is the autonomous vehicle in <Picture 2>, preserving its exact recognizable identity, white exterior, compact rounded body, dark glass, wheels, lighting, roof-mounted autonomous sensor equipment, dimensions and silhouette. Ignore the original daylight surroundings.
summary:
[reference generation] Keep <Subject 1> dancing beside the parked <Subject 2> beneath the same cyan canopy in a single restrained 1980s performance shot, ending with <Subject 1> standing at the vehicle door ready to enter.
retention_analysis:
<Subject 1> (appears throughout [Shot 1]): fully_preserved - exact identity, facial appearance, costume, proportions, metallic hands, cap and shoes remain unchanged during simple controlled choreography.
<Subject 2> (appears throughout [Shot 1]): fully_preserved - remains stationary and visually identical to <Picture 2>, with only environmental neon reflections added.
detailed_description:
Fifteen-second photorealistic single-take performance in the exact same nightlife curb location. Authentic 1980s cyberpunk film appearance: practical cyan neon canopy, tiny magenta accents, wet pavement, dark brutalist architecture, chrome reflections, steam, blue-black shadows, soft optical bloom, anamorphic flare and subtle 35mm grain.
[Shot 1] Begin with <Subject 1> leaning naturally against the front side of <Subject 2>, exactly matching the previous ending.
The camera begins a slow clockwise orbit around the character and vehicle together. Keep both reference subjects visible for almost the entire shot.
<Subject 1> gently pushes away from the vehicle and begins a restrained, repeatable Italo-disco dance phrase designed to preserve identity: two lateral steps, one controlled shoulder roll, metallic right hand sweeps horizontally across the chest, left metallic index finger points upward, a small pivot, then two measured steps backward.
Keep limb proportions stable and movements humanly achievable. Preserve the exact uncanny friendly facial expression.
<Subject 2> remains parked in precisely the same position. Cyan reflections travel naturally over its white panels and dark glass, but the vehicle shape and sensor equipment never change.
The wet ground produces soft reflections of both subjects.
As the orbit approaches completion, <Subject 1> stops dancing, turns toward <Subject 2>, walks the short distance to the door and reaches for the opening.
END STATE / TRANSITION: <Subject 1> stands immediately beside the open door of <Subject 2>, one metallic hand resting on the door frame, body oriented toward the cabin and ready to sit. The next scene begins here.
No new vehicle, no background change, no additional dancers, no body deformation, no wardrobe change, no readable text.
overall_soundscape:
None required.
non_diegetic_music:
None generated. Music added separately.
subject_definitions:
<Subject 1> is the person in <Picture 1>, preserving their exact identity, facial features, skin tone, hairstyle, facial proportions, body proportions, pale blue uniform suit, white shirt, black tie, matching round pale blue cap, metallic robotic hands, black footwear, and distinctive cheerful uncanny expression.
<Subject 2> is the autonomous vehicle in <Picture 2>, preserving its exact recognizable identity, white exterior, body shape, proportions, dark glass, wheels, front and rear lighting, roof-mounted sensor equipment and overall silhouette. Ignore the sunny location visible in <Picture 2>.
summary:
[reference generation] Continue <Subject 1> entering <Subject 2>, closing the door and smoothly departing the same cyan curb during one continuous fifteen-second tracking shot.
retention_analysis:
<Subject 1> (appears throughout [Shot 1]): fully_preserved - face, identity, clothing, accessories, proportions and robotic hands remain unchanged while transitioning naturally from standing to seated.
<Subject 2> (appears throughout [Shot 1]): fully_preserved - vehicle identity and geometry remain constant through stationary and moving states.
detailed_description:
The target is a fifteen-second photorealistic continuous shot maintaining the exact same downtown curb, same cyan lighting, wet street, brutalist architecture, subtle magenta accents, haze, steam, anamorphic bloom, soft diffusion and 1980s film grain.
[Shot 1] Begin with <Subject 1> standing beside the already open door of <Subject 2>, one metallic hand on the upper door frame.
Without cutting, <Subject 1> smoothly lowers into the front cabin. Maintain realistic limb articulation and stable body proportions. Both metallic hands move naturally inside, followed by the legs and black shoes.
The door closes.
The camera remains outside and begins gliding parallel along the side window as <Subject 2> gently pulls away from the curb.
Through the glass, keep <Subject 1> clearly recognizable under cyan dashboard illumination. <Subject 1> looks forward and taps one metallic hand lightly against the upper leg in a restrained rhythmic pattern.
<Subject 2> smoothly merges back into the exact same wet boulevard.
Camera continues beside the car for several seconds without changing distance abruptly.
During the final four seconds, camera gradually reduces speed while <Subject 2> maintains forward motion. The car naturally moves ahead until camera settles into a rear-left three-quarter view.
END STATE / TRANSITION: stable rear-left tracking view of <Subject 2> traveling away along the familiar cyan-lit boulevard. The next clip begins directly from behind this moving vehicle.
No cuts, no teleportation, no different vehicle, no character change, no new architecture, no text.
overall_soundscape:
None required.
non_diegetic_music:
None generated.
subject_definitions:
<Subject 1> is the person in <Picture 1>, preserving their exact identity, facial features, skin tone, hairstyle, facial proportions, body proportions, pale blue uniform suit, white shirt, black tie, matching round pale blue cap, metallic robotic hands, black footwear, and distinctive cheerful uncanny expression.
<Subject 2> is the autonomous vehicle in <Picture 2>, preserving its exact recognizable identity, white exterior, compact proportions, rounded geometry, dark windows, wheels, lighting arrangement, roof-mounted autonomous sensor system and silhouette. Ignore the daytime background from <Picture 2>.
summary:
[reference generation] Follow <Subject 2> through the familiar neon downtown while <Subject 1> rides inside, ending with the vehicle entering a long cyan tunnel during one uninterrupted fifteen-second night-driving shot.
retention_analysis:
<Subject 1> (appears through vehicle glass): fully_preserved - same identity, facial appearance, costume, proportions, cap and robotic hands remain consistent.
<Subject 2> (primary visual subject): fully_preserved - exact shape, white body, dark glazing, wheels, sensor equipment and proportions remain unchanged during the entire drive.
detailed_description:
Fifteen-second photorealistic continuous tracking shot in the same cyberpunk downtown, same night and same 1980s film aesthetic: practical cyan architectural lights, minimal magenta accents, wet asphalt, concrete towers, thin steam, deep blue shadows, chrome reflections, soft diffusion, anamorphic streaks and subtle grain.
[Shot 1] Begin directly behind and slightly left of <Subject 2>, matching the rear-left three-quarter ending of the previous clip.
Camera travels at approximately the same speed and maintains a consistent following distance.
<Subject 2> drives calmly through the established downtown boulevard. Wet pavement reflects the white vehicle and repeating cyan architecture.
Sparse pedestrians remain safely on the sidewalks. One cyclist travels in a separated lane.
The road gradually curves to the right. Camera follows the same smooth arc and slowly moves closer toward the left side of the vehicle.
Through the dark side glass, briefly reveal <Subject 1> seated comfortably in the front cabin, still wearing the exact pale blue uniform and cap. <Subject 1> gives one gentle head nod and one small shoulder movement while looking forward.
Do not make <Subject 1> dominate this shot; the drive itself is the focus.
Ahead, reveal a long rectangular road tunnel built into the same downtown architecture. The tunnel entrance is illuminated by repeating cyan rectangular lights.
<Subject 2> aligns smoothly with the tunnel entrance.
END STATE / TRANSITION: centered rear view of <Subject 2> just beginning to cross into the cyan tunnel, with the repeating light geometry visible ahead. Next clip begins inside this exact tunnel.
No new environment, no speed racing, no vehicle mutation, no daylight, no captions.
subject_definitions:
<Subject 1> is the person in <Picture 1>, preserving their exact identity, facial features, skin tone, hairstyle, facial proportions, body proportions, pale blue uniform suit, white shirt, black tie, matching round pale blue cap, metallic robotic hands, black footwear, and distinctive cheerful uncanny expression.
<Subject 2> is the autonomous vehicle in <Picture 2>, preserving its exact recognizable identity, white exterior, compact proportions, rounded front geometry, black glass, wheels, lights, roof-mounted autonomous sensor assembly, and overall silhouette.
summary:
[reference generation] Follow <Subject 2> through the cyan tunnel, gradually move alongside it and reveal <Subject 1> taking both robotic hands away from the controls for a small seated disco gesture before the car reaches the tunnel exit.
retention_analysis:
<Subject 1> (appears prominently in second half of [Shot 1]): fully_preserved - exact identity, face, blue clothing, cap, tie, metallic hands and body proportions remain unchanged; only restrained arm and shoulder movement is introduced.
<Subject 2> (appears throughout [Shot 1]): fully_preserved - exact vehicle geometry, sensors, windows, body panels, wheels and color remain recognizable and constant under moving tunnel lights.
detailed_description:
The target is a fifteen-second photorealistic uninterrupted shot inside the same urban tunnel. Authentic 1980s science-fiction cinematography: repeating cyan practical light rectangles, dark concrete walls, wet pavement, occasional subtle magenta reflection, light atmospheric haze, anamorphic streaking, soft diffusion and tactile film grain.
[Shot 1] Begin directly behind <Subject 2> as it completes entry into the cyan tunnel.
Camera follows the vehicle at identical speed. Repeating cyan light bands travel rhythmically across the exact white body, dark windows and roof-mounted sensor system without changing their physical forms.
After several seconds, camera slowly moves from directly behind to the left side of <Subject 2>, arriving at a clean parallel tracking composition.
Through the side window, clearly reveal <Subject 1> in the front cabin.
Maintain exact facial identity and outfit. <Subject 1> calmly lifts both metallic hands completely away from the controls and brings them loosely to chest height.
Perform only a tiny seated Italo-disco gesture: two synchronized metallic fingertip taps in empty air, one subtle shoulder pulse and one relaxed head nod.
<Subject 2> continues perfectly straight without visible human control.
The effect should feel confident, cool and slightly humorous, never slapstick.
During the final four seconds, cyan and magenta city lights become visible beyond the tunnel exit. Camera gradually advances into a front-left side position.
END STATE / TRANSITION: front-left side tracking view of <Subject 2> precisely at the tunnel exit, with the familiar nighttime city visible immediately beyond. Next scene continues the same forward movement.
No visual transformation, no speed jump, no different vehicle, no costume changes, no text.
overall_soundscape:
None required.
non_diegetic_music:
None generated.
subject_definitions:
<Subject 1> is the person in <Picture 1>, preserving their exact identity, facial features, skin tone, hairstyle, facial proportions, body proportions, pale blue uniform suit, white shirt, black tie, matching round pale blue cap, metallic robotic hands, black footwear, and distinctive cheerful uncanny expression.
<Subject 2> is the autonomous vehicle in <Picture 2>, preserving its exact recognizable vehicle identity, white body, dimensions, rounded design, dark glass, wheel geometry, front lighting, roof-mounted sensor equipment and silhouette. Ignore the original daylight environment.
summary:
[reference generation] Continue <Subject 2> exiting the cyan tunnel and traveling through the familiar downtown while <Subject 1> performs a slightly more energetic seated disco gesture, ending with the car stopped beneath the previously established cyan canopy.
retention_analysis:
<Subject 1> (appears prominently through windshield): fully_preserved - face, identity, body proportions, clothing, cap, tie and metallic hands remain exact; only controlled rhythmic motion changes.
<Subject 2> (appears throughout [Shot 1]): fully_preserved - vehicle identity, body, windows, wheels, sensor system and lighting remain unchanged.
detailed_description:
Fifteen-second photorealistic continuous shot, same downtown, same night, same weather and exact same 1980s cyberpunk color grade: cyan practical lights, restrained magenta highlights, wet asphalt, brutalist facades, steam, deep shadows, analog diffusion, anamorphic lens streaks and 35mm grain.
[Shot 1] Begin as <Subject 2> exits the cyan tunnel from the front-left side composition established previously.
Camera smoothly transitions into a low front-left three-quarter tracking position while moving backward at identical speed.
The vehicle's white surface reflects long cyan lines and occasional magenta highlights from the same familiar architecture.
Through the windshield, <Subject 1> is clearly visible and slightly more animated than before while remaining physically stable.
<Subject 1> performs two gentle shoulder pulses, one head nod and then raises one metallic hand for a playful forward finger point.
The autonomous vehicle continues operating smoothly and safely.
Camera gradually gets closer to the windshield while maintaining enough visible vehicle body to preserve <Subject 2>'s identity.
<Subject 1> slowly turns toward the camera and shows the same recognizable friendly uncanny smile.
During the last four seconds, <Subject 2> slows and returns to the exact same curb beneath the cyan canopy used earlier.
END STATE / TRANSITION: <Subject 2> completely stopped beneath the familiar canopy, viewed from a stable front-side position, with <Subject 1> visible through the window looking toward camera. The next scene begins from this exact setup.
No alternate neighborhood, no new car, no daylight, no character redesign, no text.
subject_definitions:
<Subject 1> is the person in <Picture 1>, preserving their exact identity, facial features, skin tone, hairstyle, facial proportions, body proportions, pale blue uniform suit, white shirt, black tie, matching round pale blue cap, metallic robotic hands, black footwear, and distinctive cheerful uncanny expression.
<Subject 2> is the autonomous vehicle in <Picture 2>, preserving its exact recognizable vehicle identity, white exterior, body design, proportions, windows, wheels, lighting and roof-mounted autonomous sensor assembly. Ignore the sunny environment in <Picture 2>.
summary:
[reference generation] Have <Subject 1> step out of the stopped <Subject 2> beneath the familiar cyan canopy, perform one final simple Italo-disco dance, then return to the vehicle in one coherent continuous fifteen-second performance shot.
retention_analysis:
<Subject 1> (appears throughout [Shot 1]): fully_preserved - exact facial identity, pale blue uniform, cap, tie, metallic hands, footwear and proportions remain unchanged during restrained choreography.
<Subject 2> (appears throughout [Shot 1]): fully_preserved - remains parked and visually identical to the reference while serving as a stable visual anchor.
detailed_description:
The target is a fifteen-second photorealistic continuous performance shot in the exact same cyan-lit curb location established earlier. Same wet pavement, same concrete facades, same practical cyan lighting, same subtle magenta accents, atmospheric steam, deep blue-black shadows, anamorphic flares, diffusion and textured 1980s film grain.
[Shot 1] Begin with <Subject 2> stopped beneath the cyan canopy and <Subject 1> visible through the side window.
The vehicle door opens smoothly.
<Subject 1> steps out naturally and stands beside the exact car.
Camera begins slowly pulling backward as <Subject 1> walks two measured steps toward lens.
<Subject 2> must remain clearly visible behind <Subject 1> throughout the performance.
<Subject 1> performs the final restrained Italo-disco phrase: two side steps, one metallic right-hand finger point, one controlled shoulder roll, a small half-turn, one smooth backward glide, then both metallic hands briefly rise symmetrically at chest height.
Keep the choreography simple, physically believable and identity-preserving.
Cyan light reflects across the pale blue suit and chrome robotic hands while magenta remains only a secondary accent.
After the short dance, <Subject 1> stops, looks over the shoulder toward <Subject 2>, turns and calmly walks back to the open vehicle door.
<Subject 1> begins lowering into the seat.
END STATE / TRANSITION: <Subject 1> is halfway seated inside <Subject 2>, one black shoe still on the wet pavement, door open, cyan canopy overhead. Next scene begins from this exact physical pose.
No additional dancers, no crowd, no costume change, no vehicle variation, no text.
subject_definitions:
<Subject 1> is the person in <Picture 1>, preserving their exact identity, facial features, skin tone, hairstyle, facial proportions, body proportions, pale blue uniform suit, white shirt, black tie, matching round pale blue cap, metallic robotic hands, black footwear, and distinctive cheerful uncanny expression.
<Subject 2> is the autonomous vehicle in <Picture 2>, preserving its exact recognizable vehicle identity, white exterior, compact proportions, rounded body shape, dark windows, wheels, lighting, roof-mounted autonomous sensor equipment and overall silhouette. Ignore all daylight environmental information from <Picture 2>.
summary:
[reference generation] Complete the video with <Subject 1> entering <Subject 2>, giving one small final gesture through the window, and the exact vehicle driving away through the familiar neon boulevard during one continuous eleven-second closing shot.
retention_analysis:
<Subject 1> (appears during first half of [Shot 1]): fully_preserved - exact face, identity, clothing, cap, tie, metallic hands, footwear and proportions remain unchanged during the final seating and farewell gesture.
<Subject 2> (appears throughout [Shot 1]): fully_preserved - exact white autonomous vehicle, body design, windows, wheels, sensor equipment, lights and proportions remain stable until it disappears naturally into the city.
detailed_description:
The target is an eleven-second photorealistic cinematic final single take in the exact same cyberpunk downtown district on the same night. Maintain the established 1980s science-fiction film aesthetic: practical cyan neon, restrained magenta accents, wet black boulevard, brutalist concrete buildings, chrome reflections, drifting steam, deep blue-black shadows, soft optical diffusion, anamorphic horizontal flares and subtle 35mm grain.
[Shot 1] Begin exactly with <Subject 1> halfway seated inside <Subject 2> beneath the familiar cyan canopy, with one black shoe still outside.
<Subject 1> smoothly brings the remaining leg and metallic hands into the cabin, settles into the seat and closes the vehicle door.
Through the side window, <Subject 1> turns toward camera one final time and performs a tiny understated farewell: two metallic fingers rise briefly in a restrained disco gesture.
<Subject 2> gently begins moving away from the curb.
Camera remains stationary at street level at first, watching the exact vehicle move deeper down the same familiar wet boulevard.
After several seconds, camera begins a slow cinematic crane upward, revealing the same cyan-lit brutalist architecture already established throughout the video. Do not introduce any new landmark or district.
<Subject 2> becomes progressively smaller while its white body and roof-mounted sensors remain recognizable under the neon light.
Cyan reflections stretch along the wet road behind it.
During the final seconds, <Subject 2> reaches the same distant corner previously seen in the video and turns gently behind a building.
The vehicle disappears naturally from sight.
Hold very briefly on the empty wet boulevard, cyan neon reflecting across the pavement and a small cloud of steam drifting through frame.
Slow cinematic fade to black.
No new subjects, no new vehicles, no location change, no transformation, no text, no subtitles, no generated logos.
overall_soundscape:
None required. Visual generation only.
non_diegetic_music:
None generated. Final song continues underneath during editing and fades with the image.example 9
https://www.reddit.com/r/comfyui/comments/1vj3t9h/almost_as_bad_as_the_star_wars_holiday_special/
Standard live-action The Big Bang Theory sitcom look: practical television photography style, a sitcom apartment set, basic lens, average depth of field, tv quality recording look, tv studio lighting, standard living room props, sit and stand around acting, with very little walking.
Scene overview: the apartment set from The Big Bang Theory <Picture 2>, the protagonist Sheldon Cooper <Picture 0> sitting on the couch, Luke Skywalker <Picture 1> is sitting left of Sheldon, Sheldon Cooper <Picture 0> complains to Luke <Picture 1>
Storyboard: (each shot is a wide and medium shots of the same set, cuts only when a new character is shown):
[0s-10s] Shot 1: medium shot of Sheldon <Picture 0> and Luke Skywalker <Picture 1> on the couch: Sheldon <Picture 0> is complaining to Luke <Picture 1>. Sheldon Cooper: "I don't know what to tell you Mark, the writing, th-the acting, it was just appauling. It was almost as bad as the Star Wars holiday special." Followed by a laugh track.
Camera: each shot its focused on the stage and actors, always facing the set like it would on any sitcom.
Audio: Tv studio quality. Straight from the Seinfeld tv show.
No text, subtitles, logos or watermarks of any kind, no animation or cartoon rendering, no overly-CG look, keep the classic 90s sitcom live-action texture. example 10 - CHARACTER SWAP V2V
https://app.notion.com/p/CHARACTER-SWAP-V2V-PROMPTS-halo-christo-3b663c2a7548802a8143d3ac8cb23373
https://www.reddit.com/r/comfyui/comments/1vinc36/testing_character_swap_with_minimax_h3/
SCENE CONTEXT
Object replacement pass. In <Video_1>, the target object is replaced by the object shown in <Image_1>. Everything else in <Video_1> remains exactly as it is.
ACTIVE REFERENCES
<Video_1>: the master plate. Camera path, framing, timing, cast, environment, lighting and every other object 100% match <Video_1>.
<Image_1>: identity of the replacement object only. Its shape, proportion, material, colour, logos and surface markings 100% match <Image_1>, kept legible and correctly oriented throughout. NO MASK
MOTION INHERITANCE
The replacement object inherits the full behaviour of the object it replaces, frame by frame: same screen position, same scale, same rotation, same motion path, same speed, same entry and exit timing. Whatever the original object did, the new object does identically. No new movement is introduced and none is removed.
INTEGRATION
Contact reads physically: hands wrap the new silhouette, supporting surfaces meet its actual base, contact shadows land directly beneath it, and any grip conforms to its real geometry.
Occlusion order is preserved: whatever passed in front of the original object passes in front of the new one, and whatever it covered stays covered.
Reflections, refractions and cast shadows on nearby surfaces are rebuilt for the new geometry while keeping the same direction and softness as the plate.
OPTICS
Shot size, FOV, depth of field, focus falloff and motion blur carried over from <Video_1> with no drift. The object sits at the same focal plane as the original.
CAMERA
Camera behaviour, height, distance, movement and handheld character identical to <Video_1>.
PHYSICS
Mass, inertia, swing and settle behaviour consistent with the material shown in <Image_1>. Any fluid, spill, dust or particle interaction updates to the new geometry while obeying the same gravity and timing as the plate.
LIGHTING
Key direction, intensity, falloff and white balance taken from <Video_1>. The object catches the same key from the same side, sits at the same ambient level, and throws a shadow matching the existing shadows in length, direction and softness. Specular highlights appear only where the plate's key light would place them, reading the true surface finish from <Image_1>.
STYLE
Photoreal, fully integrated into the original plate: same grain structure, same black level, same tonal contrast, same colour grade as <Video_1>.
POSITIVE LOCKS
- Only the target object changes; every other element of <Video_1> stays untouched.
- The object stays present, complete and correctly scaled in every frame the original appeared in.
- Identity from <Image_1> holds steady across the whole clip, with no drift in shape, colour or markings.
- Edges blend seamlessly: matching noise, matching edge softness, no halo, no outline.
- One continuous plate, cuts only where <Video_1> already cuts.example 11 - It’s made from 12 separate clips, 15 seconds each, so 3 minutes total. Not audio driven.
https://www.reddit.com/r/comfyui/comments/1vi2ole/testing_of_3minute_ai_music_video_rtx_3090_with/
A surreal psychedelic cinematic music-video shot. A single glossy electric-blue mouth floats in total darkness, emerging slowly from black as if waking up. Highly detailed wet luminous lips with ultraviolet highlights and liquid-cyan reflections.
The mouth starts almost still and then moves as if singing to an unheard electronic song. Simulate musical rhythm visually: at regular intervals the lips pulse open and closed and each pulse emits a circular cyan soundwave traveling outward through the darkness. Tiny glowing particles shake and react as the waves pass.
Small abstract wireless and Bluetooth-like symbols occasionally flicker around the mouth and dissolve into blue vapor.
Very slow cinematic macro push toward the lips, shallow depth of field, halation, bloom, subtle analog-video texture.
During the final seconds the mouth opens wider, revealing an intense tunnel of blue light inside. The camera accelerates forward and enters the glowing mouth.
End while the camera is traveling into the mouth so the next scene can continue from inside it.
Continue directly from inside the glowing electric-blue mouth. The camera flies forward through a surreal organic tunnel made from glossy lip textures, translucent crystal teeth, liquid membranes and rippling cyan light.
Simulate an electronic musical rhythm visually. Bright pulses travel down the tunnel at repeating intervals. Concentric soundwave rings, spirals and glowing frequency ribbons race along the walls. Tiny wireless symbols drift through the space like bioluminescent insects.
The tunnel stretches, contracts and breathes rhythmically.
The camera moves rapidly forward with smooth banking turns and gentle rolls, maintaining constant momentum.
Near the end the tunnel expands into a vast black void. A gigantic electric-blue mouth floats ahead.
As the camera approaches, that single mouth begins dividing into many identical mouths.
End at the moment the duplication begins.
Begin with the giant blue mouth from the previous scene dividing into dozens of identical glossy electric-blue mouths floating in deep black space like a constellation.
Each mouth behaves like a different musical instrument. Some pulse slowly like bass, others open and close rapidly. Each emits a different visible waveform: circular cyan rings, spirals, jagged crystalline waves, vibrating ribbons and translucent frequency walls.
Where the waves intersect they briefly create luminous geometric flowers, abstract faces and interference patterns.
The camera moves continuously through the constellation in zero gravity, weaving between mouths and passing extremely close to some of them before pulling rapidly into wide views.
Toward the end all mouths suddenly rotate toward the same distant point.
They open together and release one enormous synchronized beam of blue frequency light.
The camera accelerates along the beam.
End as the beam reaches a distant humanoid silhouette.
Begin with the cyan frequency beam from the previous scene reaching an androgynous humanoid figure standing inside a dark surreal nightclub.
The room contains reflective black floors, giant mirrors, ultraviolet fog and isolated cyan lights.
The figure initially has a perfectly smooth face with no mouth.
The incoming beam hits the face and a glossy electric-blue mouth forms there from liquid neon.
Immediately the new mouth begins moving rhythmically as if singing.
Every implied beat produces a powerful visible bass wave expanding through the room. Curtains ripple, mirrors bend elastically, reflections distort and pools of light pulse across the floor.
Camera continuously orbits around the figure, alternating between medium shots and extreme close-ups of the mouth.
The movement becomes increasingly energetic.
During the final seconds a huge soundwave hits a mirror behind the figure. The mirror liquefies into chrome water.
The camera follows the wave directly through the liquid mirror.
End while crossing through it.
Continue through the liquid mirror into an inverted psychedelic chrome-and-blue world.
Two giant glossy electric-blue mouths float facing one another from opposite sides of the frame.
They move rapidly toward each other, trailing long ribbons of cyan waveform light.
They stop only millimeters apart without touching.
Between them the air vibrates violently with visible frequency patterns.
The mouths pulse and move as if harmonizing to an unheard electronic song. Every implied beat compresses the light between them further.
A brilliant sphere of cyan energy gradually forms between the lips, made from waveform lines, liquid light and tiny wireless symbols.
The camera performs a fast continuous orbit around the mouths while moving closer.
At the climax the sphere bursts outward into thousands of miniature glowing blue mouths and fragments.
The camera immediately chooses one tiny falling mouth and dives downward after it.
End while rapidly following the falling mouth.
Follow the miniature blue mouth falling directly out of the previous scene.
It drops through darkness and suddenly enters a surreal nighttime city.
Thousands of tiny glowing electric-blue mouths are falling from the sky like rain.
The camera descends rapidly toward street level and begins flying forward between buildings.
Each mouth opens and closes rhythmically while falling. Whenever one strikes a window, car, rooftop or puddle it generates a glowing circular soundwave.
Streetlights flash rhythmically. Building reflections ripple. Neon signs distort. Puddles produce synchronized frequency patterns.
The camera races low over the wet street, occasionally passing through clouds of falling mouths and narrowly avoiding chrome objects.
As the sequence progresses, the pavement begins turning into electric-blue liquid.
Buildings stretch downward into their own reflections.
The entire city melts into one enormous blue ocean.
End with the camera racing just above the newly formed liquid surface.
Begin directly above the infinite electric-blue liquid ocean created from the melting city.
The camera flies extremely low over the surface at high speed.
Huge glossy blue mouths rise explosively from beneath the water like surreal islands.
Each mouth opens on an implied beat and sends enormous circular waves across the ocean.
The camera banks around the waves, dives between giant lips and skims across the liquid surface while glowing droplets explode upward.
A chrome humanoid figure suddenly appears running across the water.
Every footstep generates luminous frequency rings.
The camera follows beside and behind the figure with aggressive tracking movement.
Ahead, an enormous mouth rises from the ocean and opens.
The water begins rotating into a giant whirlpool shaped like an abstract Bluetooth symbol.
The chrome figure accelerates toward it and is pulled into the vortex.
The camera dives in immediately behind the figure.
End while falling into the whirlpool.
Continue falling through the blue whirlpool.
The chrome figure suddenly splits into two mirrored humanoid bodies suspended in an enormous electric-blue void.
Each body has a glowing blue mouth embedded in the chest where the heart would be.
The bodies orbit each other rapidly.
Their chest-mouths repeatedly open and fire thin cyan waveform lines toward one another.
The connection fails several times, causing sharp visual glitches, duplicate frame echoes and bursts of distorted space.
Simulate rhythmic music visually through repeated connection attempts, pulsing light and synchronized body movement.
The camera revolves around both figures while constantly changing distance, moving from extreme close-ups of the heart-mouths to wide shots of the orbiting bodies.
Finally one bright continuous waveform successfully connects both mouths.
Their movements instantly synchronize.
The cyan connection becomes brighter and thicker until it fills the center of the frame.
The camera accelerates directly into the connecting waveform.
End inside the bright line.
The bright connecting waveform from the previous scene expands around the camera and transforms into an extreme macro view of two electric-blue mouths approaching one another.
The lips are constructed from glossy liquid, pixels, tiny waveform fragments and floating droplets.
The camera moves rapidly around and between them while maintaining extreme macro detail.
As the mouths approach, streams of luminous data begin moving between them.
Tiny landscapes, abstract memories, faces, wireless symbols and pulses of cyan light flow from one mouth into the other.
The mouths never physically collide. Instead they exchange increasingly intense streams of information through the narrow gap between them.
Each implied musical beat causes the lips to pulse, the data stream to surge and the camera to change speed.
The two mouths gradually lose their solid form and dissolve into one enormous shared waveform.
The waveform twists violently into a spiral.
As the camera follows it, the spiral begins forming the petals of a giant electric-blue flower.
End as the flower starts opening.
Continue from the waveform flower opening in the previous scene.
Reveal an enormous psychedelic flower floating in black space, constructed entirely from hundreds of glossy electric-blue mouths.
Every petal is a mouth.
The flower rotates rapidly while different rings of mouths open sequentially, creating visible waves traveling around its circumference.
Each pulse generates cyan frequency ribbons, glowing pollen and expanding geometric mandalas.
The camera spirals aggressively through the flower, diving between petals, rotating around the center and constantly shifting between huge wide shots and extreme macro mouth close-ups.
The rotational speed steadily increases.
The mouth petals stretch into liquid ribbons and the entire flower transforms into an enormous rotating blue mandala.
At the climax every mouth closes simultaneously.
The entire structure collapses rapidly inward.
Everything disappears except one single isolated electric-blue mouth floating in darkness.
The camera brakes suddenly and stops close to it.
End on the lonely mouth.
Begin with the isolated electric-blue mouth floating alone in nearly total darkness.
The atmosphere is quieter but still constantly moving.
The camera slowly circles the mouth while drifting closer and farther away.
The mouth attempts to sing.
Each time it opens, a visible cyan waveform travels outward but quickly breaks apart into digital particles.
The mouth tries repeatedly.
Every failed signal creates analog glitches, transparent duplicates and brief distortions of the surrounding darkness.
The pulses become weaker.
Then a tiny blue signal appears extremely far away.
The camera suddenly changes direction and moves toward the distant light while keeping the mouth visible behind.
The isolated mouth reacts and fires a stronger waveform.
The distant signal answers with another wave.
Both waves race toward each other through the void.
The camera accelerates alongside them.
End one instant before the two signals collide.
Begin exactly before the two cyan signals collide.
They connect and immediately produce an enormous radiant blue shockwave.
The camera is thrown backward at extreme speed.
The entire psychedelic universe from the previous scenes suddenly appears around the camera: hundreds of electric-blue mouths, liquid oceans, chrome bodies, waveform flowers, falling mouth rain, rotating wireless symbols and glowing frequency ribbons.
Everything moves in synchronized rhythmic pulses.
The camera races through the environment, passing between giant mouths, diving through waveform rings, rolling around chrome bodies and skimming above liquid-blue surfaces.
Every implied beat triggers another transformation.
The movement reaches maximum intensity.
Then the camera begins an extremely fast continuous pull backward.
As it moves away, the entire impossible universe becomes smaller.
Eventually reveal that everything we have been seeing exists only as a reflection on the surface of one enormous pair of glossy electric-blue lips floating in black space.
The camera continues pulling back.
The giant lips slowly close.
One final cyan soundwave explodes outward toward the camera, expanding until it fills the entire frame.
The light fades rapidly to pure black.
Hold black for the final second.example 12
https://www.reddit.com/r/comfyui/comments/1vgze70/minimax_h3_image_to_video_my_example/
Warner Bros cartoon, Wile E. Coyote run back of Beep Beep. Beep beep, running along and doesn't realize where he's going. So Beep Beep ends up crashing into a tree. Then the scene change with Wile E. Coyote is cooking on the barbecue with a single cooked Beep Beep on. He puts it in his mouth, chews it, but makes a disgusted and then spits it out to the side. Then scene change with Wile E. Coyote seen from back is entering in a Mc Donald's restaurant.example 13
https://www.reddit.com/r/StableDiffusion/comments/1vgqvtu/first_video_attempt_yeah_minmax/
The dwarf from <Picture 1> stands in his original environment. He warms up with a broad, friendly smile, looks directly at the camera, and raises his hand in greeting. He speaks in a friendly but gruff Scottish accent. He finishes speaking, raises his foam topped tankard in greeting, it sloshes around sloppily and drips down the side of the mug. He gives a playful wink to the camera, takes a large drink of the beer, foam soaking into his facial hair realistically. He lowers the mug and lets out a content sigh as holds his smile as the video ends.
Timeline:
[0s-1s] The dwarf brightens up, smiles warmly, and raises his hand in greeting.
[2s-10s] He speaks directly to the audience in a gruff Scottish accent: "Alright, So I've been messing around with that A.I. stuff again. McCoy, you seeing this shit? I was a dragon age screenshot once"
[10s-12s] He gives a quick, cheerful wink to the camera and raises his sloshing foam topped tankard as if in a minor toast.
[12s-15s] He takes a large drink of the beer, the foam soaking into to his facial hair where appropriate.
[15s-18s] He holds his warm smile as the scene smoothly settles to a close.
Audio: Clear spoken male voice with a distinct Scottish accent, soft movement rustle during the wave, and subtle ambient room tone.
example 14 - Prompt For Multi Character
Used multiple reference images for each scene.
https://www.reddit.com/r/comfyui/comments/1vgf9ai/assemble_the_multiverse_minimax_h3_r2v_is_awesome/
subject_definitions:
<Subject 1> is [CHARACTER 1] from <Picture 1>, preserving the visual identity, costume, proportions, accessories, and visual style shown in the reference image.
<Subject 2> is [CHARACTER 2] from <Picture 2>, preserving the visual identity, costume, proportions, accessories, and visual style shown in the reference image.
<Subject 3> is [CHARACTER 3] from <Picture 3>, preserving the visual identity, costume, proportions, accessories, and visual style shown in the reference image.
summary:
[reference generation] A 5-second cinematic multiverse portal arrival. Three characters emerge from a consistent amber-orange portal and take a calm, confident formation.
retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - retains the identity, costume, proportions, accessories, and visual style shown in <Picture 1>.
<Subject 2> (appears in [Shot 1]): fully_preserved - retains the identity, costume, proportions, accessories, and visual style shown in <Picture 2>.
<Subject 3> (appears in [Shot 1]): fully_preserved - retains the identity, costume, proportions, accessories, and visual style shown in <Picture 3>.
detailed_description:
A 5-second cinematic portal-arrival scene at dusk. One stable medium three-shot, framed from the knees up. No dialogue, no combat, no wide landscape, no camera movement, and no crowd.
Portal continuity: a large circular amber-orange portal stands behind the characters. It has a bright rotating outer ring, a darker transparent centre, floating gold sparks, gentle smoke, and a low resonant hum. Keep its colour, size, brightness, particle density, and rotation speed consistent across every portal-arrival clip.
[Shot 1] <Subject 1> steps through the portal first and takes the centre position with quiet confidence. <Subject 2> emerges on one side, naturally adjusts or lowers any item they are carrying if applicable, then gives a focused glance toward the unseen distance. <Subject 3> walks through last, takes position on the opposite side, and calmly surveys the scene. The three hold a poised, united stance as the portal flickers and golden particles drift around them. Their expressions and body language remain confident and appropriate to their individual character identities.
overall_soundscape:
Low portal hum, soft wind, subtle movement from clothing or equipment, and drifting golden sparks. No speech.
non_diegetic_music:
A restrained cinematic rise builds across the shot and resolves on a calm, confident note.
Prompt For single characters:
subject_definitions:
<Subject 1> is [CHARACTER] from <Picture 1>, preserving the visual identity, costume, proportions, accessories, and visual style shown in the reference image.
summary:
[reference generation] A 5-second cinematic multiverse portal arrival. One character walks through a consistent amber-orange portal, then takes a confident action stance with a subtle grin.
retention_analysis:
<Subject 1> (appears in [Shot 1] and [Shot 2]): fully_preserved - retains the identity, costume, proportions, accessories, and visual style shown in <Picture 1>.
detailed_description:
A 5-second cinematic portal-arrival scene at dusk. No dialogue, no crowd, no wide landscape, and no combat.
Portal continuity: a large circular amber-orange portal stands behind <Subject 1>. It has a bright fiery rotating outer ring, a darker transparent centre, floating gold sparks, gentle smoke, and a low resonant hum. Keep its colour, size, brightness, particle density, and rotation speed consistent across every portal-arrival clip.
[Shot 1] Medium knee-up shot. <Subject 1> walks steadily through the portal toward the camera, then comes to a composed stop. Their costume, silhouette, movement style, and any character-specific accessories remain fully consistent with <Picture 1>. Golden sparks drift around them as the portal flickers behind.
[Shot 2] Close-up of <Subject 1>. They shift into a distinctive, character-appropriate action stance, looking directly ahead with calm confidence. Their expression changes into a subtle smile and restrained grin. Keep the movement natural and controlled, with no exaggerated facial distortion. The portal remains softly visible and out of focus in the background.
overall_soundscape:
Low portal hum, soft wind, subtle movement from clothing or equipment, and drifting golden sparks. No speech.
non_diegetic_music:
A restrained cinematic rise builds through the entrance and resolves as <Subject 1> holds the final stance.example 15 - using a small black image as first frame
tip - writing a prompt and inputting an image as the last frame. However, the default workflow seemed to always require a first frame. So, I used a small black image as the first frame; this time, it worked.
https://www.reddit.com/r/StableDiffusion/comments/1vl3uu4/minimax_h3_testing_l2va_moon_landing/
integrated_multimodal_description: Time-lapsed, cinematic, a medium-wide shot of a film set in a large indoor studio which is used to shoot a scene of moon landing involving lunar module, US flag and an astronaut. At 00:00.000 the camera shows an empty, sterile white studio room, with recognizable vertical wand in the back and horizontal floor at its bottom. At 00:01.000 Some film crews install a black wand into the studio's vertical wand. The black wand has some tiny white shining points which represent stars. At 00:02.000 Some workers fill the studio's floor with some dirty-white sand, gravels and small rocks and form a barren lunar landscape. At 00:03.000 Some film crews bring an Apollo Lunar Module and place it into the left side of the scene. At 00:04.000 A film crew places a US flag with pole on the right side of the scene. Another film crew puts a picture of the Earth as the blue planet, partially blacked on its bottom side, on the top right corner of the scene. At 00:05:000 An astronaut walks in into the scene, goes into the middle of the scene, faces to viewer and waves his hand. At 00:07:000 The whole scene settles into the exact arrangement, position of subject and objects, camera angle, lighting, and final composition established by <Picture 1>. A male deep voice of the director (S1) says, <d>[English] Cut!</d>
overall_soundscape:
non_diegetic_music: Sustained violin notes at a very fast tempo with spaced piano tones.
reddit posts with lots of general info
https://www.reddit.com/r/comfyui/comments/1vinc36/testing_character_swap_with_minimax_h3/
https://www.reddit.com/r/StableDiffusion/comments/1vhuorq/45_lower_minimax_h3_sampler_time_with_new/
https://www.reddit.com/r/comfyui/comments/1vknr0v/followup_from_a_5second_clip_to_a_247/
https://www.reddit.com/r/NeuralCinema/comments/1vk20gc/minimax_h3_hack_50_reference_or_more/
https://www.reddit.com/r/comfyui/comments/1vjs1re/is_there_a_nice_place_to_learn_how_to_use_minimax/
https://www.reddit.com/r/comfyui/comments/1vjn8i9/minimaxh3_speedup_nodes_is_this_connected/
https://fal.ai/learn/devs/minimax-h3-prompting-guide
https://www.reddit.com/r/comfyui/comments/1vj8nyp/forcing_minimax_h3_to_generate_multispeaker/
tricks and tips
about audio:::
https://www.reddit.com/r/comfyui/comments/1vjbyqk/minimax_is_nutso/
I’ve had better luck not just saying “no voices,” but explicitly building the prompt around non-vocal audio.
Something like:
“No dialogue, no voiceover, no singing, no speaking, and no vocalizing characters. The scene is entirely atmospheric.”
Then list the sounds you actually want, like wind, footsteps, traffic, rain, doors, clothing rustle, etc.
For example:
“overall_soundscape: Soft night ambience, distant traffic, light wind, footsteps on pavement, car unlock chirp, door handle click, and the muted thump of the car door closing.”
And if you don’t want background music either: “non_diegetic_music: N/A”
Also don’t give the character an (S1) speaker ID if they never talk. That seems to help keep the model from inventing speech.