I update frequently. As soon as I see a bug, I kill it. I apologize for the crazy morning of chaos updates!
2.1.2 — reference images, per-shot saves, and six repairs
One new capability and a set of fixes. The chaining itself is untouched: nothing here changes how shots join, so a graph that renders well today renders identically after updating.
New: reference images actually have a way in
The sampler's reference_images input has existed since 2.0, and SETTINGS.md documented it — as unwired, because nothing in the shipped workflow was connected to it. There was no way to use it without building the lane yourself.
The full workflow now carries a REFERENCE lane in the anchors column: two image loaders → ImageBatch → a gate → the sampler. It ships with the gate off, so nothing changes until you turn it on. Point the loaders at portraits of your character, flip the gate, and they ride into every shot as <Picture 1>, <Picture 2>. Chain another ImageBatch for a third and fourth. Needs a ref2va checkpoint.
Unlike the identity anchor above it, these are not a first frame — they do not constrain shot 1's composition, they only carry who the person is. That also makes them the thing that covers shot 1, where the memory bank is still empty and has nothing to carry identity from yet.
New: every shot can be saved as it renders
A chain only became a file at the very end, when the master was muxed. So anything that failed after the last shot destroyed the entire run — and one report was three hours lost to an out-of-memory error at the mux, after every shot had already rendered successfully. The work existed and was thrown away, which is the worst version of that bug rather than the mildest.
save_every_shot (on both samplers) writes each shot to output/video/H3_SHOTS/ the moment it decodes, alongside the master. If the mux dies you have every shot on disk and a joining job instead of a lost day. Files are written before the seam trim, so consecutive shots overlap by about a second — the master is still the clean join, these are the safety copy. Requested in issue #13.
New: custom sigma schedules
The samplers built the schedule themselves from steps + scheduler, with no way to supply your own — so a turbo LoRA that ships the curve it needs could not use it, and ran wrong rather than refusing. Both samplers now take an optional SIGMAS input. Connect one and it replaces the schedule entirely, steps rebinds to len(sigmas)-1 so the two-pass upscale split is taken as a fraction of your curve, and the console states that the steps and scheduler widgets are being ignored instead of silently overriding you. Link-only input, so saved graphs load unchanged. Issue #14.
Fixed: --- separators were ignored in passthrough mode
example_script.txt ships --- separated and every document tells you to write scripts that way, but the writer's passthrough path returned the whole file as ONE shot — which the sampler then repeated to fill shot_count. Pasting a finished four-shot script rendered the entire text as shot 1, four times over. It now splits on exactly the same rule the sampler uses. A single paragraph is still one shot, so .txt batches are unaffected.
Fixed: a stale prompt-set filename killed the whole queue
ComfyUI validates every combo value in a graph before it runs anything. If RiftPromptSource's saved source_file no longer existed — a renamed folder, a workflow shared from another machine, or the prompt lane simply switched to manual — the queue failed with Value not in list and nothing ran, including the lanes that were perfectly fine. Switching to manual did not disable it, because validation happens before the switch is ever consulted.
The node now declares VALIDATE_INPUTS, so the filename is resolved only if the node actually executes. Manual mode now genuinely disables it. If it does run and the file is missing, the error names the file.
Fixed: the writer
These three only matter if you let the LLM write your shots. If you paste your own shot list, they change nothing for you.
Every story came out 15 shots. The system prompt ordered exactly 15 whenever the brief didn't ask for a count, so the model never got to decide. It now counts the story's beats and lands where the story lands — measured 4–7 on ordinary briefs, and about 7 when the brief gives no length signal at all. Seven chained shots is roughly 65 seconds at 243 frames, which clears the 1-minute mark platforms pay on. Ask for a specific count and you still get exactly that count.
The mode dropdown did nothing. The shipped workflow's system_prompt box held a frozen copy of the long-story prompt, and a filled box overrides the per-mode prompt file — so every mode ran the long prompt no matter what the dropdown said, and prompt updates in the pack never reached anyone using the shipped graph. The box now ships empty and the dropdown works. If you saved your own copy of the v2.0 or v2.1 workflow, clear that box by hand — your saved graph still carries the old frozen prompt, and this update cannot reach it. Related: short_story is now 1–3 shots instead of always exactly 1.
A messy answer from the model no longer kills the render. Three separate real failures, all seen on local models: a markdown fence sharing a line with the payload (```json {"prompts":…) was destroyed by the fence stripper; a reply truncated at the token ceiling left no balanced object to recover; and parsing sat outside the retry loop, so one malformed answer ended a run after a 100–240 s call that a re-ask usually fixes. The order is now clean → parse → retry ×3 → salvage the shots that completed → fail, and the final error names passthrough mode as the escape hatch.
Verification, plainly. The reference lane, the per-shot saves and the --- fix are all confirmed by a live chained render, not by inspection: the console logs 2 reference image(s) ride in every shot as <Picture 1..2>, PASSTHROUGH: 4 shot(s) (it would have said 1 before the fix), and shot 1/4 saved through shot 4/4 saved, with four files on disk. The three JSON failures are unit-tested against the actual captured payloads that caused them; the live re-run afterwards parsed on the first attempt, so that proved no regression rather than proving the salvage path fires in a real queue. Shot counts were measured over two runs of five briefs on a local 8B: 4–7 shots, ±1 between runs — before the change all five returned 15. The custom sigmas input has not been exercised by a render — it needs a schedule source wired in, and I would rather say so than imply it.
Install
Unzip. Copy both node folders into ComfyUI/custom_nodes/:
ComfyUI-H3-Multishot/ the sampler and helper nodes
ComfyUI_JoyAI_Echo_GGUF_Nodes/ the LLM prompt writer (full workflow only)
Restart ComfyUI. ComfyUI v0.30.0 or newer is required — that is the release with native MiniMax-H3 support.
Load a workflow from workflows/ through the workflow menu.
The full workflow needs five packs (two are in this zip): the writer pack below, ComfyUI-H3-Motion-Context for the context_pin default, RES4LYF for the beta57 scheduler, and ComfyUI-sol-attn + comfyui-minimax-h3-blockcache-T8 for the VRAM panel; ComfyUI-Custom-Scripts adds the script preview. ComfyUI validates every node class before it will queue, so a missing pack stops the whole workflow rather than just its own feature — INSTALL.md lists how to remove each one instead if you would rather not install it. CORE needs none of them, and that is tested on a clean install.
On the Motion-Context fork. There is an active fork, ethanfel/ComfyUI-MiniMaxH3-Contex-Loop, well ahead of upstream with a disk-backed chain/loop system. It is a complement, not a replacement: by design it leaves the MiniMaxH3MotionContext node id to NikoDemon80's pack. Install the fork instead of upstream and context_pin still fails, because the node it calls is not there. Install both - they are built to coexist, and this pack works with either one's runtime patches. As of 2.1.1 the error message says so directly when it detects the fork.
The writer pack is RealRebelAI's (github.com/RealRebelAI/ComfyUI_JoyAI_Echo_GGUF_Nodes), modified so the workflow's join rules actually reach the model; NOTICE_RIFT_MODIFICATIONS.md inside it lists every change. If you already have that pack, replace it with this copy. The CORE workflow does not need it at all.
Models you need
checkpoint MiniMax-H3 ref2va (GGUF Q8_0 / Q5_1 / Q4_0) -> models/diffusion_models
text encoder qwen3vl minimax_h3 (+ its -mmproj sidecar) -> models/text_encoders
video VAE minimax_h3_video_vae -> models/vae
audio VAE minimax_h3_audio_vae -> models/vae
GGUF quants: huggingface.co/joeygambino/MiniMax-H3-GGUF — Q8_0 for 32 GB, Q5_1 for 24 GB, Q4_0 below that.
GGUF encoder pairing. ComfyUI-GGUF matches the -mmproj vision sidecar to the encoder by filename, in the encoder's own folder. Rename either, or split them up, and it loads the encoder without its vision tower — which presents as the model ignoring your reference image. This pack's CLIP loader raises instead of continuing blind, uses the only mmproj beside the encoder when there is exactly one, and takes an mmproj_name widget so you can point at the file directly.
Which workflow
H3_Seamless_Chain_CORE — start here. The same seamless chaining with zero third-party packs. Type shots into the script box and queue.
H3_Seamless_Chain_v2 — everything: master controls, LLM writer, VRAM panel, identity and voice anchors, episode/batch prompt source, boundary plates, audio spine. Optional lanes are gated off by default.
H3_Keyframes — one clip, anchors at chosen frame positions, per-anchor condition strength.
The two things that stop people on the first run
1. The prompt writer needs a model you have pulled
The full workflow points at a local Ollama with model_name = qwen3:14b. If it is not pulled, the first queue stops immediately:
LLM API error 404: model 'qwen3:14b' not found
Fix: ollama pull qwen3:14b. Any OpenAI-compatible endpoint works — its URL in base_url, its exact tag in model_name. ollama list prints the tags you have, and it must match character for character.
No LLM at all? Set the master panel's use_file_prompts to manual entry, delete the writer, and feed your own script into the sampler's script input — one prompt per shot, separated by --- on its own line. CORE already works this way.
2. A local writer will fight the video model for the card
Turn on unload_model_after on the writer. It frees its own model from Ollama the moment the script is written. Without it the model stays resident for the server's default five minutes — your whole first shot. ComfyUI's own eviction cannot reach it, because Ollama is a separate process with its own allocator, and Ollama's OpenAI-compatible endpoint has no keep_alive field to ask with; the switch calls the native endpoint, which honours it. On under 32 GB, prefer a remote endpoint entirely.
Settings: start here, change nothing
checkpoint ref2va sampler euler
continuity context_pin scheduler beta57 (full) / beta (CORE)
steps 14 fps 24
frames/shot 362 (~15.1s, the trained maximum)
resolution 1280x736 landscape or 768x1344 vertical
beta57 comes from RES4LYF, not stock ComfyUI. Measured on an identical seed it scored 10/10 for lip-sync against 8/10 for stock beta, with image quality, skin texture, artifacts and audio judged equal — so the full workflow ships it and lists RES4LYF as required, while CORE ships beta and keeps its zero-third-party-pack promise. One widget either way.
Leave every VRAM switch off and the reserve at 0, and try a render before touching any of it. The activation reserve measures each shape and conditioning payload as it renders and sizes the pool itself; it holds on 24 GB cards as well as 32 GB. A hand-set reserve overrides that measurement, so a number that suited one shape becomes wrong for the next. Those switches exist to dig out of a spill the console has already named, not for pre-emptive tuning.
Resolution cannot change mid-chain, and the mux must stay at 24 fps — other rates audibly shift voice accents. Dial-by-dial reference in SETTINGS.md.
Writing a script that chains cleanly
The previous shot's last ~1 second is replayed at the head of the next and discarded. Four rules follow, and breaking them is what produces mid-word chops and pose jumps:
Open holding. Every shot after the first opens in the previous shot's exact closing arrangement, with no dialogue for ~2 seconds. Give it real micro-motion — a breath, a weight shift — so it does not read as a freeze.
Land settled. Every shot ends with ~2 seconds of quiet, back in a stable arrangement, all dialogue finished.
Never split a line across shots. Dialogue plus 4 seconds of hold and settle must fit the shot length. If it does not fit, move the whole line to the next shot.
Repeat descriptions word-for-word. Character appearance and the room/light description, byte-identical in every shot. An unnamed light source gets reinvented per shot, and that is where colour drift starts.
The LLM writer applies these for you. Hand-written scripts must follow them — PROMPTING.md has a worked four-shot example, and example_script.txt is ready to paste.
How the chaining works
context_pin carries the previous shot's last 22 frames as raw latents — never decoded to pixels and re-encoded — placed at interior keyframe coordinates, with a timeline-placed audio reference alongside. The regenerated head is trimmed on decode. Colour, motion and voice cross the boundary as data rather than as a description.
Motion is the clearest case. Hand the next shot a single frame and it knows position but not velocity, so pace can reset at the boundary. Measured on a steady-pace walk: a single-frame anchor with no memory bank wobbled at the join; context_pin held it, and so did the memory bank on its own.
first_frame is the alternative — the model's own trained hand-off, no extra pack, and what CORE ships with. cut for episodic work.
Identity and voice
Nothing wired — the frame relay plus verbatim descriptions hold a face surprisingly well. A ~40 s two-character scene held both faces with no reference images at all.
self_anchor_voice (on) — shot 1's own rendered voice becomes the reference for every later shot. No file needed; write shot 1 with a clean solo line.
voice_ref — a clean solo speech clip, pinned across the whole chain including shot 1.
reference_images — character portraits carried into every shot as <Picture 1..N>. Bind them in the prompt text.
seed_per_shot (leave on) — measured: varying the seed per shot holds the face; one seed for every shot drifted both face and voice. Identity lives in the conditioning, not the seed.
When something goes wrong
404, model not found — the writer's model is not pulled. See above.
A word clips at a join — the script put dialogue too close to a boundary. Move the whole line; do not split it.
Sharpening increases every shot — the texture ratchet. Set chain_gain_control to flatten; worth it past about 5 shots.
Stalls at 0 steps, or runs several times slower than usual — a VRAM spill, the driver paging to system RAM instead of erroring. The console now names it. Raise the reserve, or drop resolution, frames, or reference payload.
Red or missing nodes — an optional pack is not installed. Delete those nodes, or use CORE.
GGUF architecture error — the pack teaches ComfyUI-GGUF the minimax_h3 architecture at startup. If it persists, run python apply_gguf_arch_patch.py from the pack folder once and restart.
Audio dulls on a very long chain — expected; restart the chain on a scene cut, where a fresh start costs nothing.
Fixed in 2.1.1
context_pin died when ComfyUI-H3-Motion-Context was installed. Both packs patched the same ComfyUI method and Motion-Context refuses to stack on an unrecognised wrapper, so its payload patch failed and the chain errored. This pack's wrapper now does everything theirs does and declares their compatibility marker, so whichever loads first owns the site and the other stands down. Load order no longer matters. Verified with a live context_pin render.
seamless_tail crashed mid-chain with Motion-Context installed ("only first/last keyframe anchors are supported") - after your first shot had already rendered. It needs interior keyframe anchors, which conflict with that pack; it now stops before any sampling with the alternatives named: use context_pin, or first_frame, or remove that pack.
seamless often reads as a cut - now labeled. It is a legacy latent-only soft pin kept for comparison; the model satisfies it loosely. For a real join use context_pin or first_frame. The tooltip and settings reference now say so plainly.
Why these never showed in testing here: an install-layout difference disabled the conflict detection on the dev machine. That detection is fixed, and release testing now runs on a packaged clean install so the class cannot slip through again.
Everything you need is in the zip. One download: both node folders, all three workflows, and the full documentation. Nothing else to fetch, nobody to ask.
What changed in 2.1
The full workflow now works on a clean install. It referenced a prompt-source node that had never been published, and drove the writer through inputs the upstream writer pack does not have — so the boundary rules never reached the model. It rendered, and it rendered worse than it should, with no error to explain why. Both fixed: that node ships here as RiftPromptSource, and the rules are written into the workflow's own system prompt as well as carried by the writer pack in this zip.
The chaining sampler's anchor switches now do something. voice_ref, reference_images, self_anchor_voice, preview_first_shot, two_pass_upscale, reference_image_size and the sampler/scheduler overrides were drawn on the canvas but absent from the class, so ComfyUI stripped them before execution. All real now, render-verified.
unload_model_after on the writer, described above.
SHOT COUNT on the master panel drives the sampler and the writer together so they cannot disagree; prompt source switches between a manual scene box and a prompt set, lazily.
Node titles no longer name a checkpoint or a switch position — a title like H3 model (fl2va) is a lie the moment you change the model.
Two-pass upscale cannot be combined with context_pin or latent_handoff, or with an audio spine: those carry raw latents, or one locked denoise trajectory, across the join, and a two-pass render preserves neither. The node stops with an error naming the conflict rather than quietly producing a weaker join. Two-pass is available on cut, seamless, seamless_tail, first_frame and flf_chain.
Credits
Prompt writer: RealRebelAI (ComfyUI_JoyAI_Echo_GGUF_Nodes, modified — see the NOTICE in the zip). context_pin: NikoDemon80 (ComfyUI-H3-Motion-Context). Two-pass upscaling: Tr1dae (ComfyUI-MiniMaxH3_LatentUpscaler).