Hi, I need help with a character-replacement / dubbing setup on Muse Minimax Director (v1.2) using 3 references at once:
- Ref 1 (Picture): person A identity image
- Ref 2 (Picture): person B identity image
- Ref Video: source video containing two people whose performance I want to copy
What I want:
A shot-for-shot dubbing of the reference video — the two persons from the pictures must reproduce the video's exact body movement, hand gestures, head motion and lip/mouth movements (frame by frame), while camera, framing, wardrobe, props, location, lighting and audio are all copied from the reference video. The pictures should contribute identity only.
My setup:
- Mode: Reference (Omni)
- Total Duration 15s / Chunk Size 15s (single chunk)
- use_prompt_override: ON (custom text pasted into the Prompt Override socket)
- Ref Video: Role = Reference, Subject description left blank (pure motion/camera reference), "Include this clip's audio" = ON
- Retention: fully preserved on all references
- Ref Image Size: max
- Dialogue language: Arabic
Result so far:
- ✅ Face/identity replacement works — the two new faces appear and stay stable
- ❌ Motion is NOT copied from the reference video — the generated video invents its own movement instead of reproducing the source performance (body, hands, and mouth movements don't match the reference video)
What I tried:
- My override text originally referenced only <Picture 1>/<Picture 2> and described the video in plain words ("the reference video") without a <Video N> tag.
- I was told the issue might be that the prompt never contains an explicit <Video 1> tag, so the model receives the video but isn't instructed to follow its motion.
My questions:
1. In Prompt Override mode, how exactly should a wired Ref Video be referenced in the text? Is <Video 1> the correct tag, and how is it numbered when only some ref slots are filled (dense numbering like <Subject N>, or fixed to the slot number)?
2. Does the compiled prompt normally inject the <Video N> tag automatically, and is it expected that Prompt Override bypasses that injection — meaning I must write it myself?
3. Is "Subject description = blank" the right way to get pure motion/camera reference, or does a blank description actually cause the video to be ignored/downweighted?
4. Is there any known limitation of the Reference (Omni) model regarding motion transfer from ref video (vs. identity transfer from ref images)?
5. Any recommended retention combination for this exact use case (dubbing/motion cloning)?
Environment:
- Workflow: muse_minimax_h3_director V1.4
- Mode: Reference (Omni), single chunk, 15s
- VRAM: 31.84 GB, Maximum Quality profile
- Sampler: euler / beta, steps 10, two-stage sampling ON
Any guidance on the correct prompt-override syntax for ref video would be greatly appreciated. Thanks!