
Prototype a controlled presenter shot from three visual references and supplied speech, then run an RTX 2× upscale stage.
Included: One editable multi-reference graph with turbo sampling and RTX 2× postprocessing, setup notes and input checklist.
Requirements: Current ComfyUI, MiniMax H3 reference model and nodes, Qwen3-VL encoder, video/audio VAEs, turbo LoRA, RTXVideoSuperResolution node, compatible NVIDIA GPU, and your licensed inputs. Weights, nodes and media are not included.
Before buying: this is an editable ComfyUI workflow template, not rendered footage. Higher complexity and memory demand; the customer graph and upscale stage have not yet been rerun together.