MiniMax-H3 Multishot — Seamless Chain: multi-shot scenes that render as one continuous take (picture + audio)
184
7.5k
148
Download
1 variant available
4610 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 9
(19)
Aug 7, 2026
MiniMax H3
If you ever downloaded the H3 GGUF quants and got an error the instant ComfyUI touched the file, this is the fix — and the error was never your fault or the quant's.
GGUF models now load with no manual step
ComfyUI-GGUF validates a GGUF's architecture against a fixed list and rejects anything not on it before reading a single tensor. Upstream's list has no minimax_h3 entry, so every MiniMax-H3 DiT quant died with:
Unexpected architecture type in GGUF file: 'minimax_h3'
The pack has always shipped apply_gguf_arch_patch.py to fix that, but it was one line buried in the install steps. In practice people hit the error, concluded the models were broken, and gave up — three separate people reported it across the model and workflow repos, and those are just the ones who said something.
The pack now does it for you. On startup it adds the architecture to ComfyUI-GGUF's live list in memory. Nothing is written to disk, it is idempotent, and unlike the on-disk patch it survives ComfyUI-GGUF updates instead of being reverted by them. You will see this in the console:
[H3] taught ComfyUI-GGUF the 'minimax_h3' architecture
The old script is still in the folder as a fallback for unusual setups, but you should not need it. If you previously ran it, nothing breaks — the new code sees the architecture is already known and does nothing.
The other GGUF error: text encoder and the mmproj file
Different problem, same week, so it is worth spelling out here. If your H3 text encoder GGUF fails with a state_dict or vision mismatch against its -mmproj file, load it with this pack's H3 Clip Loader (Any) rather than the stock CLIPLoaderGGUF.
The H3 encoder is a truncated Qwen3-VL-32B — 50 layers, no final norm, no lm_head — and its vision tower ships separately as the -mmproj-F16.gguf sidecar. Stock ComfyUI-GGUF only merges an mmproj when the encoder's architecture is qwen2vl. Qwen3-VL reports qwen3vl, so the sidecar is never merged at all, and the missing vision tensors surface as a state_dict mismatch. Its mmproj key map is qwen2vl-era besides: wrong merger keys, and no rules for H3's deepstack mergers or split QKV.
This pack's loader does all three things stock cannot — truncates the text tower, merges the sidecar explicitly, and renames the vision tensors to H3's layout. Keep the -mmproj file in the same folder as the encoder and do not rename either one: they are paired by filename.
And to answer the question that came up directly: a full .safetensors encoder (fp8, int8, NVFP4-AWQ, whatever) works without any of this because it is a complete, pre-shaped model with the vision tower already inside. That is a property of the container, not of the quantization — NVFP4 is not doing anything special.
Nothing else changed
Same nodes, same workflows, same defaults as v1.4. Existing graphs render identically. Upgrading is just overwriting the folder and restarting.
Thanks to the people who took the time to report the error instead of quietly writing the models off. That is the only reason it got fixed.
Show more

1710 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 9
5450 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 9
15.9K0 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 9.0 1 2 3 4 5 6 7 8 9K

MiniMax H3 is licensed by MiniMax under the MiniMax H3 Community License Agreement. That agreement’s Applicable Territory excludes the European Union, the United Kingdom, the Republic of Korea and the United States of America. Your use of H3 and of any H3 derivative is subject to that agreement and its Acceptable Use Policy.
MiniMax H3
Support
Everything here is free and stays free — the format spec, the nodes, the workflows, the cartridges, the LoRAs. If it saved you a night of debugging (it contains several hundred of mine), tips keep the 5090 warm:
🔁 Liberapay (recurring)
⚡ Or right here: the Civitai tip button on this page sends Buzz directly.
Type a story. Get one continuous video, with sound. Multi-shot scenes render as a single take - no last-frame chaining, no quality loss from shot to shot. That's the whole pitch.
Both workflows now read left to right: numbered lanes, and you only ever touch lanes 2-4. Everything below the main row is optional.
What you need
ComfyUI + this node pack (Manager: MiniMax-H3 Multishot, or the zip on this version).
A MiniMax-H3 checkpoint (links on this page). 24 GB card? Take a GGUF.
SEAMLESS CHAIN - multi-shot scenes as one take
Lane by lane:
README - the quick start lives on the canvas itself.
1 - MODELS - pick your H3 checkpoint and text encoder. VAEs are preset; LoRA slots are empty until you fill one.
2 - ANCHORS (optional) - a photo to open shot 1 on (enable its gate), and a short voice clip to lock the speaker's voice.
3 - YOUR PROMPTS - type your idea in the box, or point the switch at a prompt file. The writer expands it into shot prompts. Writing your own? Set the writer to
passthrough (raw JSON, skip LLM)and paste shots separated by---lines.4 - CONTROLS - size, frames per shot, steps, and
take_seconds(total length; 30 is a good first run). The switches stay off unless you installed the pack a switch names.5 - ENGINE - nothing to change. The remote encoder lives here if you want the text encoder on a second PC: enter its address, flip the encoder switch, free ~15 GB.
6 - OUTPUT - your video and its audio save here.
Optional panels below the main row: reference images (your character, ref2va checkpoints - folder per character + AUTO REFS on), V2V reference (a clip whose look guides the render), FFLF plates (flf_chain mode only), audio spine (a soundtrack the take follows).
EXTEND TAKE - one person talking, as long as you want
Same lanes, different job: one premise becomes ONE continuous speech cut across windows.
1 - MODELS - same as above.
2 - ANCHORS - a photo of your speaker (shot 1 opens on them) and a voice clip. More useful here than anywhere: one person carries the whole take.
3 - YOUR PROMPT - ONE premise, one speaker. The writer writes the whole speech. num_shots 0 = it decides. Passthrough works here too.
4 - CONTROLS -
take_secondsis the star: 30 ships, 60 clears TikTok's minute.windowstays on auto - it sizes itself to your card.5 - ENGINE / 6 - OUTPUT - same as above.
Keep takes to about 4 windows for now - very long takes slowly sharpen.
Rules of thumb (both workflows)
Spoken lines: 8-12 words per shot. Short lines sync; long lines garble.
Say the sounds you want ("rain on the roof, a fridge hum") or it invents its own.
Keep your character's face in frame - faces carry identity between shots.
If something breaks
Red node? Update the pack in Manager, restart, reload the workflow from disk.
Render crawls at low wattage? Lower resolution or frames per shot, or use the remote encoder.
Only one of the two workflows shows in your sidebar? Fixed in 2.6.5 - re-download both.
Still stuck: comment with your console log. I answer.
Deep dives: the two articles linked on this page. Every lane also has a short note on the canvas.
Detailed guide for people that can read good:
Every setting explained: the Seamless Chain deep manual | Civitai