Download
1 variant available
This checkpoint includes a config file, download and place it along side the checkpoint.
1,3180 1 2 3 4 5 6 7 8 9,0 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 9
(62)
Aug 7, 2026
MiniMax H3
First version release

25.7K0 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 9.0 1 2 3 4 5 6 7 8 9K
4K0 1 2 3 4 5 6 7 8 9K
75.7K0 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 9.0 1 2 3 4 5 6 7 8 9K
⚠️ This is a beta version, and more features will be added as the community comes up with useful ideas and improvements.
⚠️ Warning: Although this is not an especially advanced workflow, It uses custom nodes and, depending on the setup, may require basic knowledge of Python dependency management, virtual environments, the ComfyUI file system, and troubleshooting. If you are unfamiliar with these topics, please make a backup of your ComfyUI setup before installing new nodes or updating ComfyUI.
For beginners, I generally recommend using the portable version, as it is easier to maintain and back up. It is also important to have a solid understanding of how to properly use and prompt the MiniMax H3 model.
📝 Sage Attention nodes are not included. You can enable Sage Attention globally by adding the --use-sage-attention flag to your startup script.
💡You can find the system prompt I use to generate my MiniMax prompts here.
The Prompt Enhancer is not included because I haven't been able to find a ComfyUI-compatible model powerful enough to consistently handle the system prompt. I personally use Qwen3.6 running on Ollama.
⚠️ This workflow is built with the standard ComfUI layout system in mind, Nodes 2.0 will mess up the size of the nodes and front end functions like the bypasser switches, we are all grown ups here, let the kids play with toys 🤭
Main Features
Anything-to-video: Text-to-video, image-to-video, audio + image-to-video, video-to-video, and more.
Turbo mode switch
Upscaling and frame interpolation
Character identity lock
Lip-sync
Motion transfer:V2V / body performance
Camera movement / cinematography transfer
Voice cloning
Video editing
Multi-reference fusion: audio, video, and images
Native stereo audio generation: dialogue, SFX, and music generated jointly
First / last frame conditioning
Text and brand rendering
Video extension / continuation
Style transfer
Relighting / object swap / background replacement: as an editing subset
Multi-shot / multi-scene storytelling
Multilingual dialogue
📝 Personal Notes
The recommended video length is 20 seconds maximum. Beyond that, results can become inconsistent, although I've managed to successfully push the length up to 30 seconds on an RTX 4090.
For voice cloning, use reference audio around 5 seconds long. Longer references may cause the reference audio itself to leak into the final video.
The sample videos were generated using the LTX upscaler included in the workflow, along with a high-quality first-frame image, to achieve the best possible output quality.

