Sign In

MiniMax H3 Ref2VA | All Inputs + Turbo mode + LTX 2.3 Upscaler + NVIDIA VSR + Frame Interpolation

Download

1 variant available

Config Other

MinimaxR2VCivit.json

340.12 KB

Verified:

Type
Workflows
Stats

1,318

Reviews
Published

Aug 7, 2026

Base Model

MiniMax H3

Hash
AutoV2
7779B51023
default creator card background decoration
Reactions - 25673

25.7K

Followers - 4037

4K

Downloads - 75671

75.7K

⚠️ This is a beta version, and more features will be added as the community comes up with useful ideas and improvements.

⚠️ Warning: Although this is not an especially advanced workflow, It uses custom nodes and, depending on the setup, may require basic knowledge of Python dependency management, virtual environments, the ComfyUI file system, and troubleshooting. If you are unfamiliar with these topics, please make a backup of your ComfyUI setup before installing new nodes or updating ComfyUI.
For beginners, I generally recommend using the portable version, as it is easier to maintain and back up. It is also important to have a solid understanding of how to properly use and prompt the MiniMax H3 model.

📝 Sage Attention nodes are not included. You can enable Sage Attention globally by adding the --use-sage-attention flag to your startup script.

💡You can find the system prompt I use to generate my MiniMax prompts here.

The Prompt Enhancer is not included because I haven't been able to find a ComfyUI-compatible model powerful enough to consistently handle the system prompt. I personally use Qwen3.6 running on Ollama.

⚠️ This workflow is built with the standard ComfUI layout system in mind, Nodes 2.0 will mess up the size of the nodes and front end functions like the bypasser switches, we are all grown ups here, let the kids play with toys 🤭

Main Features

  • Anything-to-video: Text-to-video, image-to-video, audio + image-to-video, video-to-video, and more.

  • Turbo mode switch

  • Upscaling and frame interpolation

  • Character identity lock

  • Lip-sync

  • Motion transfer:V2V / body performance

  • Camera movement / cinematography transfer

  • Voice cloning

  • Video editing

  • Multi-reference fusion: audio, video, and images

  • Native stereo audio generation: dialogue, SFX, and music generated jointly

  • First / last frame conditioning

  • Text and brand rendering

  • Video extension / continuation

  • Style transfer

  • Relighting / object swap / background replacement: as an editing subset

  • Multi-shot / multi-scene storytelling

  • Multilingual dialogue

📝 Personal Notes

  • The recommended video length is 20 seconds maximum. Beyond that, results can become inconsistent, although I've managed to successfully push the length up to 30 seconds on an RTX 4090.

  • For voice cloning, use reference audio around 5 seconds long. Longer references may cause the reference audio itself to leak into the final video.

The sample videos were generated using the LTX upscaler included in the workflow, along with a high-quality first-frame image, to achieve the best possible output quality.