product image

MiniMax H3 Two-Reference Speaking Character | ComfyUI Workflow

$55

Combine a character portrait, scene reference and supplied speech to prototype a consistent speaking presenter.

Included: One editable reference-to-video graph, a reference preparation guide and a speaking-shot prompt framework.
Requirements: Current ComfyUI with MiniMax H3 reference-to-video nodes, MiniMax H3 reference model, Qwen3-VL encoder, video/audio VAEs, two images and speech supplied by the buyer. Weights and media are not included.
Before buying: this is an editable ComfyUI workflow template, not rendered footage. The customer graph has not yet been rendered; speech alignment, identity retention and duration require a real test with your inputs.