Download
2 variants available
Config Other
MiniMaxH3_LipSync_V4_N2_Captions_DiskTimeline.json
394.59 KB
Verified: a day ago
Generation, training and LoRA distribution on Civitai are covered by Civitai’s own license agreement with MiniMax. If you download these weights and run them yourself, your use is instead governed by the MiniMax H3 Community License Agreement, whose grant excludes the European Union, the United Kingdom, the Republic of Korea and the United States of America.
MiniMax H3
MiniMax H3 LipSync V4 — Captions and Disk Timeline Guide
Workflow: MiniMaxH3_LipSync_V4_N2_Captions_DiskTimeline.json · Revision: 4.0 · Updated: 2026-09-26
Generated songs or uploaded audio · full-song lyric alignment · animated captions or clean export · exact disk-backed frames · full accumulated previews · FL2VA or Ref2VA conditioning · silence-aware transitions
What Is This Workflow?
This workflow turns a generated song or uploaded audio track and one or more images into a continuous MiniMax H3 video, with optional animated lyric captions. Generate with YuE2 or MiniMax Music 3, save the full song, then choose the excerpt that H3 should perform. You can also choose Selected file in Full song source and mute unused music branches.
The excerpt defines the master audio and video duration. H3 renders manageable tasks; the workflow writes their retained frames to temporary exact-pixel chunks and streams those chunks into previews and the final render. The caption branch aligns lyrics against the full song, then applies the same excerpt start and duration as H3. Both captioned and clean exports use the selected soundtrack. The separately saved song stays full length.
This guide covers the latest DiskTimeline variant. The earlier MiniMaxH3_LipSync_V4_N2_Captions.json also has captions but uses accumulated IMAGE batches; it does not provide the disk-backed loop described here.
Use it for:
A single character singing, speaking, or performing through a long track.
Multiple images that take ownership of different parts of the timeline.
Evenly spaced or manually timed image changes.
Transitions aligned to pauses and resumed speech or singing.
Independent source-image resets, FL2VA first/last interpolation, or exact hard cuts.
Long same-image sections that either continue visually, reset plainly, or turn a necessary generation split into an intentional camera cut.
FL2VA positional keyframes or Ref2VA semantic identity/reference conditioning.
Word or line captions, including floating and rotating styles, with an optional clean-video fallback.
The base preview shows the first completed task. The accumulated preview grows after every extension, so you can inspect the result before the full timeline is finished.
Recommended first test: use a short audio excerpt, one image, FL2VA, warmup_reset, a two-second warmup, and native 0.999 visual conditioning strength.
What Changed in Revision 4.0?
Full song source and H3 clip (shared caption trim) separate complete song analysis from the excerpt used for lip-sync. Shared clip start drives both the H3 excerpt and caption trim; captions use the excerpt's actual duration.
H3 Lyric Captions adds full-song alignment, six caption styles, optional full-song vocal analysis, an alignment report, clip-relative SRT, and reusable full-song timing JSON. Manual lyrics can override the selected generator lyrics.
Enable lyric captions selects caption rendering or a clean H3 fallback. The final Save Video with Generation Data node accepts either route.
Decode and Append Timeline replaces growing IMAGE concatenation. Each completed task adds an exact temporary chunk; the loop carries small references instead of the complete decoded video. Generation settings and native VAE decoding are preserved.
Preview Video still shows every completed scene in order. The preview MP4 is a viewing copy; continuity and final rendering read the exact chunks.
Trim Frame Timeline selects the final range without rebuilding an IMAGE batch. Frame Timeline Loop Gate also handles plans needing only the base task.
Long runs trade accumulated frame RAM for temporary disk space: about 29.5 GiB for four minutes at 512 × 896, 24 FPS, RGB float32, plus previews and exports. See Temporary Disk Space and Cleanup below before a long run.
The YuE2/Music 3 branches, full-song FLAC saving, H3 planning rules, seeds, sampling precision, and separate soundtrack routing remain available.
4.0 is the workflow revision; 4.4.10 is the required Eclipse package version. The three Eclipse H3 planning/conditioning nodes retain their actual V2 names:
MiniMax H3 Audio Timeline Planner V2 builds the complete frame-accurate task plan and a tensor-free analysis manifest.
MiniMax H3 Audio Plan Step V2 resolves each task's audio, original image, optional generated continuity, guides, legal render length, crop, and retained frame count.
MiniMax H3 Segmented Conditioning V2 creates either FL2VA positional keyframes or one persistent Ref2VA active-image reference. It never combines both conditioning families.
The older transition_mode, bridge_span, bridge_window_seconds, and cropped_lookahead_conditioning controls do not belong to V2. Their replacements are segment_strategy, warmup_seconds, conditioning_family, technical_split_source, and technical_seam_style.
V2 also makes hidden warmups, task ownership, prompt ownership, guide positions, technical splits, and expected image ownership explicit. No crossfade conceals a bad seam.
Manual transition times and technical-cut directions remain visible connected text inputs. One shared Visual Conditioning Strength controls both render paths. Image Resize, after Image Selector, normalizes the selected source images to the generation canvas before planning and conditioning.
Saved Revision 4.0 Settings
These are the settings in the latest saved DiskTimeline workflow, not a requirement for every song. Connected controls override local widget values.
Music source: Full song source uses
Auto; YuE2 is active and Music 3 is muted. YuE2 Auto Prompt is off; ABC planning is on.Song duration caps: YuE2
360seconds; Music 3120seconds.Lip-sync excerpt: shared start
0, requested duration225.5seconds; clamped to the available song. H3 clip usesIncoming audio.Stop (Result Review): on for H3 clip; off for Full song source.
Audio Demux: on for H3. The separate full-song caption-vocals node is also active and connected.
H3 canvas / FPS:
512 × 896/24.H3 sampling: 8 steps, CFG
1.0,euler/simple; eight-step LoRA enabled.Timeline: FL2VA,
warmup_reset,3.0-second warmup,first_only.Technical seams:
original_image_reset+intentional_camera_cut.Visual Conditioning Strength:
0.9; use0.999for the neutral first test.Maximum render length:
362frames per task.Captions: enabled; manual lyrics muted;
en, deviceauto,rotating-words, Turning sign, Quicksand Bold at60px, glow on, caption transparency6%, unmatched passages set totranscribe.Caption canvas / FPS: linked to the final H3 timeline and shared FPS, so the effective size is
512 × 896at24FPS.Final export: MP4/H.264, CRF
19, presetveryfast; workflow and generation metadata enabled. The clean route usesshortestaudio/video duration trimming.
The shorter, neutral presets later in this guide deliberately use a two-second warmup and plain resets instead of this saved editorial-cut configuration.
Before You Run It
Show the Native 0.999 Value Correctly
To enter and see 0.999 exactly, open ComfyUI Settings, go to LiteGraph → Node Widget, enable Disable default float widget rounding., and set Float widget rounding decimal places [0 = auto]. to 3. Reload the page after changing either setting, as required by ComfyUI. These settings only control Float-widget entry and display precision; they do not add noise to the image or change H3 conditioning by themselves.
Update ComfyUI and Required Custom Nodes
Your ComfyUI installation needs the native MiniMax H3, MiniMax Music 3, and YuE2 nodes used by the saved graph. Update ComfyUI if these nodes are missing, even when you intend to leave one music branch muted.
Install or update these custom-node packages:
ComfyUI Eclipse 4.4.10 or newer — disk-backed frame timelines, streaming captions/previews/exports, full-song lyric alignment, audio review, V2 planning and conditioning.
ComfyUI Smart Model Loader — H3 and music-model loaders, shared seed/sampling settings, and the LoRA stack.
ComfyUI SmartLLM — optional lyrics/style assistance and the Music 3 / YuE2 response-splitting nodes.
ComfyMath — task-index arithmetic inside the render subgraphs.
ComfyUI-sol-attn — scheduled H3 attention, fused modulation, and feed-forward chunking.
ComfyUI-DiffAid-Patches — the experimental MiniMax H3 text-activation patch in the dedicated DIFF-AID group.
ComfyUI-Easy-Use — the extension loop.
Audio Separation — optional Hybrid Demucs vocal extraction.
Restart ComfyUI after updating, then load the downloaded workflow again.
If the timeline, caption, or H3 V2 nodes are missing, update Eclipse to 4.4.10 or newer and restart the backend before reloading the workflow. Reloading the browser alone does not register new Python nodes.
Required Models
Place the files in their matching ComfyUI model folders and reselect them in the loaders. H3 needs its video model, text encoder, VAEs, and activity encoder; music weights are only needed for the generator you enable. Muted branches still require their node packages to open the complete workflow.
Default FL2VA diffusion model — ComfyUI/models/diffusion_models/
Optional Ref2VA diffusion model — ComfyUI/models/diffusion_models/
Text encoder — ComfyUI/models/text_encoders/
VAEs — ComfyUI/models/vae/
Activity encoder — ComfyUI/models/audio_encoders/
Choose the encoder that matches the spoken language:
Chinese/base: wav2vec2-chinese-base_fp16.safetensors
English/large: wav2vec2_large_english_fp16.safetensors
The saved workflow selects the English/large encoder. This encoder only finds quiet gaps and resumed activity; it never replaces either audio stream.
Additional Models Used by Revision 4.0
Enabled H3 eight-step LoRA — ComfyUI/models/loras/
The saved selector uses a minimax/ subfolder, strength 1.0, and eight H3 sampling steps. Select the path where you actually placed the file.
YuE2 — ComfyUI/models/checkpoints/
The YuE2 Smart Model Loader uses this checkpoint with its baked text encoder and VAE. No separate Music 3 encoder or DAV belongs on the YuE2 path.
MiniMax Music 3 — separate diffusion model, text encoder, and DAV
FP32 diffusion model: minimax_music3_dit_fp32.safetensors →
ComfyUI/models/diffusion_models/.BF16 text encoder: minimax_music3_text_encoder_pruned_bf16.safetensors →
ComfyUI/models/text_encoders/.Audio VAE / DAV: minimax_music3_dav.safetensors →
ComfyUI/models/vae/.
The saved workflow uses local filenames minimaxMusic_v30_fp32.safetensors and minimaxMusic_v30_txt_3107729.safetensors. Select your downloaded Music 3 files in the corresponding loader fields. Those local filenames identify the saved choices; they are not a checksum guarantee of identity with the linked files.
Optional SmartLLM model — Ollama backend storage
huihui_ai/qwen3.5-abliterated:9b-Claude, selected in Smart LM Loader as
huihui_ai-qwen3.5-abliterated-9b-Claude-Ollama.
Both music branches use this language model for lyrics/style assistance. Configure the Ollama backend in SmartLLM, or disable that branch's Auto Prompt control and fill its manual inputs. This model belongs to the backend's model storage, not the ComfyUI H3 or Music 3 text-encoder folder.
Optional vocal separation — Torchaudio Hybrid Demucs
HDEMUCS_HIGH_MUSDB_PLUS supplies
hdemucs_high_trained.ptthrough Torchaudio's automatic download/cache.
Audio Separation manages this checkpoint in the Torch Hub cache, rather than ComfyUI/models/diffusion_models/. It is needed when Audio Demux is enabled. The separate Optional full-song vocals node uses it for caption analysis.
Caption alignment — local Whisper large-v3
Install Eclipse's requirements.txt in the Python environment that runs ComfyUI. Caption alignment uses stable-ts==2.19.1 and faster-whisper>=1.2.1,<2; font validation uses fonttools. Automatic alignment requires the five files model.bin, config.json, tokenizer.json, vocabulary.json, and preprocessor_config.json from Systran/faster-whisper-large-v3, pinned revision edaa852ec7e145841d8ffdb056a99866b5f0a478. Place them in ComfyUI/models/whisper/large-v3/. Eclipse verifies these artifacts; it does not automatically download them during execution. Valid corrected timing JSON bypasses alignment inference.
Caption font — ComfyUI/models/fonts/
The saved workflow selects Quicksand-Bold.ttf. Select an installed .ttf or .otf covering the song's script. Empty font folders receive Eclipse's bundled defaults; Pillow needs RAQM support for shaping and right-to-left layout. These caption requirements are separate from H3's Qwen and Wav2Vec models.
Match Sampling Steps to the Acceleration LoRA
The saved FL2VA setup enables the eight-step LoRA at strength 1.0, with eight steps, CFG 1.0, euler, and simple. The stack remains optional. If you change the H3 acceleration LoRA, use the step count and other sampler settings recommended for that exact LoRA. Different acceleration LoRAs target different schedules, so a step count copied from another LoRA can reduce motion, prompt response, or image fidelity.
If you disable acceleration, switch to sampling settings appropriate for the base checkpoint. Keep the selected LoRA and step count fixed when comparing conditioning or transition controls.
Select Your Inputs
Before queuing, choose your music generator or fallback audio file, select your excerpt and timeline images, verify their order, and enter one visual prompt or one prompt per image. Then confirm that the checkpoint family, text encoder, VAEs, optional LoRAs, sampler, scheduler, step count, output size, and output location match your installation.
Keep the workflow at 24 FPS, and keep generation width and height divisible by 32. Recheck all file and folder selectors before sharing a workflow or embedded metadata.
Quick Start
Choose the song in Full song source (generated or selected file). For generated music, use
Autoand enable exactly one generator. For an existing recording, chooseSelected fileand mute both generators. Leave this full-song loader at start0, duration0.Set Shared clip start (seconds) and a short
durationin H3 clip (shared caption trim). Leave its source onIncoming audio. With Stop (Result Review) enabled, queue and audition the excerpt. Queue unchanged again to continue, or disable the stop when ready.Leave Audio Demux enabled for vocal-stem conditioning, or disable it to use the master mix directly.
Select one or more timeline images and confirm their order.
Enter one visual prompt, or one non-empty prompt per image.
For the recommended neutral first test, use:
conditioning_family = fl2va_keyframessegment_strategy = warmup_resetwarmup_seconds = 2.0reset_anchor = first_onlytechnical_split_source = original_image_resettechnical_seam_style = plain_resetmax_render_frames = 362
Leave Manual Transition Times disconnected or blank for even ownership, or connect one scalar string with one comma-separated time per image after the first.
Keep Visual Conditioning Strength at
0.999for native H3 compatibility.Match the sampling step count to the selected acceleration LoRA, or use base-checkpoint settings when no acceleration LoRA is active.
Enable activity-gap alignment if transitions should move toward nearby quiet gaps and resumed activity.
For captions, supply matching full-song lyrics and verify the language, font, and local Whisper model. For a clean first test, mute the caption renderer using its H3 Render Lyrics control in Enable lyric captions.
Check free space on the ComfyUI temporary drive as well as the output drive.
Queue the video run, inspect the generated H3 Plan Report and first-task preview, then check the accumulated preview as extensions complete.
If captions are enabled, review the alignment report and captioned result after the final H3 task. Save useful timing JSON separately for later edits.
Final rendering follows the completed H3 timeline and its exact planned trim. A plan needing no extensions goes straight from the base result to final output.
A long timeline can require many H3 generations. Test image order, prompts, checkpoint family, and transition behavior on a short excerpt before committing to a full song or dialogue track.
Model Patch Groups and Fallbacks
The workflow keeps optional model modifications in separate groups. The final selector prefers the most downstream active branch and falls back through the patch chain to the loader/LoRA model when groups are muted.
The saved workflow enables scheduled attention, feed-forward chunking, and Diff-Aid; Fused Modulation is disabled. The final model selector prefers MODEL_H3_DIFF, then MODEL_H3_SOL, then MODEL_H3_LORA.
Scheduled attention, fused modulation, feed-forward chunking, and Diff-Aid are independent experiments, not universal quality presets. Compare optional patches with identical inputs and seed. To establish a clean baseline, mute the complete optional group rather than bypassing one node while leaving its branch selected.
Audio and Prompt Setup
Master Audio
V4 uses two Load Audio [Eclipse] nodes:
Full song source (generated or selected file) provides the complete recording for lyric alignment. Keep start
0and duration0here.H3 clip (shared caption trim) receives that full recording and selects the excerpt defining
MASTER_AUDIO_H3. Its start comes from Shared clip start; its duration output supplies the actual, possibly clamped caption duration.
The excerpt determines H3's duration and the clean export soundtrack. The caption renderer receives the full song and applies the same trim to its soundtrack. H3-generated audio is discarded; neither export substitutes a Demucs vocal stem.
Music generator → full-song FLAC save
↓
Full song source ← selected file option
├→ full-song caption alignment / optional full-song vocals
└→ H3 clip → MASTER_AUDIO_H3 → planner, H3 vocals, previews, clean save
Shared clip start → H3 clip start + caption trim_start
H3 clip duration output → caption duration
Final exact H3 timeline → caption background or clean save
Keep every H3 duration/planning consumer on the excerpt's trimmed output. Caption alignment is the deliberate full-song branch. Do not feed the already trimmed excerpt into caption audio while also applying the shared start: that would shift the song twice. The caption background starts at clip frame zero.
Choose Generated Audio or a File
Enable exactly one music group. The saved workflow enables YuE2 and mutes MiniMax Music 3. If both are enabled, the source selector prefers YuE2, but both full-song save output nodes may run.
In Full song source, Auto prefers valid incoming AUDIO and falls back to the selected file when no audio arrives. Selected file explicitly chooses the file even with an incoming cable connected; Incoming audio requires that input and has no file fallback. Malformed incoming AUDIO raises an error.
To use a file, choose Selected file here and mute both generators so their independent full-song save nodes do not request unwanted generation. Keep H3 clip on Incoming audio, and use matching manual full-song lyrics for captions.
Generate with YuE2
The saved YuE2 branch has Auto Prompt off and uses its manual style/lyrics. Enable Auto Prompt to turn a song concept into style and lyrics through SmartLLM's YuE2 Music task and Split YuE2. Enable ABC planning is on; native YuE2 ABC generation supplies the music generator's ABC input. Turning it off selects the empty-ABC path.
The saved duration cap is 360 seconds, with 32 diffusion sampling steps, CFG 1.0, dpm_2 / sgm_uniform, and regular audio VAE decoding. Smart Model Loader supplies the linked seed to ABC generation, music generation, and the sampler; connected inputs override the native nodes' displayed seed widgets. The full decoded song is saved as FLAC before excerpt selection.
Generate with MiniMax Music 3
Mute YuE2 and enable Music 3. Its saved Auto Prompt control is off; enable it for SmartLLM's Song Lyrics → MiniMax Music 3 assistance and Split MiniMax Music 3, or enter the manual caption and lyrics. Native MiniMaxMusic3TextEncode supplies conditioning and the duration used by the empty audio latent.
The saved duration cap is 120 seconds. The actual KSampler uses 30 steps, CFG 1.5, euler / simple, and denoise 1.0; the encoder uses CFG 1.5 and top-k 50. The loader's linked seed drives both encoder and sampler. The saved decode switch selects tiled audio VAE decoding. The full result is saved as FLAC independently of the chosen lip-sync excerpt.
Audition an Excerpt Before Lip-Sync
Set Shared clip start and H3 clip's
durationin seconds. Duration0means to the end; requests extending beyond the source are clamped. An empty excerpt or a start beyond the end is rejected. The second output reports the actual duration.Enable Stop (Result Review) and queue. Load Audio publishes its preview and interrupts execution before downstream lip-sync runs. It can execute for review without a downstream consumer.
Use the player and precise seek slider to audition. Start/duration changes update the preview without queuing; the time display is relative to the clip. Downstream AUDIO changes on the next run.
Queue unchanged again to continue from the reviewed result, or disable the stop and queue. Changed inputs require review again. Unchanged upstream results may reuse ComfyUI's cache; randomized seeds or changed inputs can regenerate the song.
The connected Shared clip start updates the audio player too. Edit that control, rather than a disconnected/local start widget, to keep H3 and captions on the same range. Keep the full-song loader untrimmed.
Incoming audio creates a temporary preview of the full source, so later excerpt edits can audition another part of it. The player identifies whether you are hearing incoming audio or the selected file, and changing the source connection clears an old incoming preview. A missing preview after restart requires another execution; it is never a persistent fallback song. The saved file selection and trim settings remain available.
Trimming preserves the source's samples, channels, batch, and sample rate without resampling. The browser preview uses 16-bit PCM and plays the first batch item; it does not change the output tensor.
Visual Performance Prompt
Describe the visible performance rather than asking H3 to compose a soundtrack:
A natural cinematic performance. The subject sings in time with the supplied
external audio, with expressive but realistic facial movement, stable identity,
coherent body motion, soft stage lighting, and controlled camera movement.
For instrumental passages, consider adding:
Keep relaxed closed lips during instrumental sections and sing only when vocals
are audible.
Visual prompting and model behavior can still produce occasional mouth motion; vocal separation is not a guaranteed lip lock.
One Prompt or One Prompt Per Image
Prompt Per Image (Blank Lines Ignored) maps non-empty lines to timeline images:
One non-empty line applies everywhere through prompt-index fallback.
With multiple lines, line 1 belongs to image 1, line 2 to image 2, and so on.
Blank lines are ignored and do not reserve an image position.
If an image has no matching line, the workflow falls back to the first prompt.
Each task receives one complete prompt. A first/last bridge from image A to image B keeps prompt A; prompt B begins with B-owned output at transition frame T. Prompts never change halfway through one generation task.
Optional Demucs Vocal Conditioning
Audio Demux controls the Hybrid Demucs branch.
When enabled:
The selected master excerpt enters Hybrid Demucs.
The raw vocal stem conditions both H3 sampling paths.
The same vocal stem feeds Wav2Vec activity analysis.
Preview Audio lets you hear the extracted stem.
When disabled, the workflow's three fallback selectors use the master audio for:
The base H3 task.
Every extension H3 task.
Wav2Vec activity analysis.
Keep Hybrid Demucs, its Preview Audio, and the conditioning-audio Set node under the shared Audio Demux control. In this saved graph, the planner's conditioning_audio comes directly from Hybrid Demucs, while sampling and activity analysis use vocal-first fallback selectors.
The master always controls duration and the final mux. A separate conditioning stem must begin at the same time and may differ in duration by no more than one 24 FPS frame. Conditioning slices keep their native sample rate and are padded only where a task boundary requires it.
This is the excerpt-only H3 stem. Caption alignment has its own full-song vocal input; the two stems have different timing origins when clip start is nonzero.
Lyric Captions and Clean Export
Choose the Lyrics for the Full Song
The lyrics selector tries Manual full-song lyrics, then active YuE2 lyrics, then active Music 3 lyrics. The generator routes use the lyrics selected for music generation. In the saved workflow the manual node is muted. Enable it and enter matching full-song lyrics to override the generator text, especially for an uploaded recording. Leave it muted when using generator lyrics.
Supply the complete sung arrangement, including repeated sections. Caption cleanup removes recognized production/section cues without changing the music generator's text. Generated singing can differ from its prompt, so review the alignment report even when lyrics came from the same branch.
Keep Alignment and Excerpt Timing Together
H3 Lyric Captions receives the full song, full-song lyrics, shared clip start, and H3 clip's actual duration. The completed, trimmed H3 timeline is its visual background. Width and height come from Final Exact Timeline Trim and FPS from the shared H3 setting. Leave those connections in place.
Optional full-song vocals is active in this saved version. It analyzes the complete recording independently of Audio Demux. You can mute this node to align against the original mix. A connected caption vocal stem must start at full-song time zero and match its duration within 50 ms. Do not connect H3's excerpt-only stem here. In both cases the original song remains the soundtrack.
Set language to the sung language, or Auto for detection; the saved value is en. Captions do not translate lyrics. Recognition of singing can be incomplete, particularly around sustained notes, backing vocals, or altered repetitions.
The saved unmatched_passages = transcribe can add recognized words in gaps where supplied lyrics could not be matched. Matched lyrics retain their wording. Review transcribed_passages in the report/timing JSON; set omit to keep only supported supplied lyrics. Neither mode guarantees that every line is found.
Choose a Caption Style
whole-line: fixed complete lines.
active-word: fixed lines with highlighting for words that have timing.
floating-words / floating-lines: moving phrases or complete lines.
rotating-words / rotating-lines: one turning text block at the selected position. Turning sign uses a vertical axis; Flipping card uses a horizontal axis. Rotation speed follows lyric timing.
The saved look uses rotating words, Quicksand Bold at 60 px, bottom-center placement, white text, a gray outline, cyan/blue glow, and 6% caption transparency. Transparency 0% is fully visible; 100% hides the caption effect. Fades are 0.5 seconds each, minimum display is 0.9 seconds, and automatic phrase grouping uses up to three words where the available timing allows it. These appearance controls do not alter H3 generation.
Review and Correct Timing
The caption node exposes:
Caption alignment report: supported and unresolved lines, analysis source, warnings, and any transcription additions.
Clip-relative SRT: subtitle times starting at the exported excerpt's zero.
Full-song timing JSON: source-song times before trimming, suitable for correction and reuse.
Cleaned lyrics: the text used for alignment.
The Show nodes display these strings; they do not automatically write subtitle or timing files. Copy useful results to separate files if you want to retain them.
To correct timing, copy the full-song JSON into Optional corrected full-song timing JSON and connect it to the renderer's corrected_timing input. That text node is blank and disconnected in the saved workflow. Keep the JSON consistent with the same full recording and cleaned lyrics; do not paste clip-relative SRT times into it. Word text and character offsets must remain consistent when editing words; use words: [] for a line-only correction.
Valid corrected JSON bypasses automatic alignment and lazy full-song vocal separation. Disconnect it to request fresh alignment. timing_adjustment is a caption-only offset: positive values delay text; it does not move the H3 video or soundtrack. Use Shared clip start to choose a different part of the song.
Appearance-only changes can reuse cached alignment and cached H3 frame chunks within the running session. Reuse depends on unchanged upstream inputs and cache retention. Restarting, clearing caches, or deleting chunks can require generation again; saved timing JSON alone cannot restore the video frames.
Export With or Without Captions
In Enable lyric captions, the H3 Render Lyrics entry controls the renderer; the separate manual-lyrics entry controls its override text. Mute the renderer to select the clean route. Use mute, not bypass.
Captioned VIDEO / clean H3 frames picks the active caption VIDEO first, otherwise the exact H3 timeline. The final Save H3 V4 Captions (or clean fallback) node is Save Video with Generation Data:
Caption VIDEO keeps its embedded soundtrack, FPS, and duration. Separate save AUDIO/FPS/duration-trim controls do not replace those values.
The clean timeline streams at its stored FPS and uses
MASTER_AUDIO_H3as its soundtrack. Duration trimming is supported; loop matching/blending is not supported for timeline inputs.
This selector saves one chosen route per run. To keep both, save the captioned result, then mute captions and queue the clean export while the H3 frames remain cached and generation inputs stay unchanged. Each uses the timestamped V4 output prefix. The base and accumulated previews remain clean H3 previews in either mode.
Timeline Images and Transition Timing
Image Order
The selected image batch is the timeline order:
One image remains the active identity/reference for the complete video.
Two or more images divide the timeline into ownership intervals.
Transition frame
Tbelongs to the destination image.
Review the thumbnails carefully. Source-folder filename sorting can change the order you expected.
Compatible identities, composition, lighting, and aspect ratios are easier to connect. Large differences are allowed, but they ask H3 to invent a more difficult handoff.
Manual Transition Times
Enter one comma-separated time in seconds for every image after the first in the connected String Multiline [Eclipse] node. Keep this as one unchanged string; the multiline node makes it easy to inspect and edit, but it does not change the comma-separated parser.
For four images:
8, 17.5, 26
This assigns image 2 near 8 seconds, image 3 near 17.5 seconds, and image 4 near 26 seconds. Values are relative to the selected excerpt, not the original full song. They must be strictly increasing and remain inside the excerpt.
Leave the text node blank, or disconnect the planner input, to divide the complete audio evenly among the selected images.
When activity alignment is enabled, even or manual times are search targets. The planner may move them to a nearby qualifying gap. Disable alignment when manual times must remain exact.
Activity-Gap Alignment
The recommended activity-alignment settings are:
align_to_activity_gap = enabledtransition_edge = activity_resumesearch_window_seconds = 5.0min_gap_duration = 0.25resume_hold_duration = 0.15
Enable alignment when timing targets may move to nearby qualifying gaps. Leave it disabled when manual transition times must remain exact.
activity_resume places the ownership change when sustained voice activity returns after a pause. silence_start places it when the pause begins.
For multiple images, search_window_seconds controls how far before and after each target the planner searches. If no qualifying gap is found, it keeps the original target.
With one image, there are no image transitions. The planner may instead use eligible quiet gaps as same-image technical split points. If none is available before a task reaches its allowed size, it splits at the maximum legal point.
Choosing the V2 Conditioning Family
FL2VA Keyframes — Default
Use:
conditioning_family = fl2va_keyframes
FL2VA gives H3 positional start and optional end images. The active original is tokenized as <Picture 1>. A real bridge tokenizes source and destination in that order.
Use FL2VA when you need:
A strong opening-image anchor.
Exact destination images at hard cuts.
First/last interpolation between two timeline images.
Optional 22-frame generated continuation at technical splits.
Keep the Smart Model Loader on matching FL2VA diffusion weights.
Ref2VA Active Reference
Use:
conditioning_family = ref2va_active_reference
Then load matching Ref2VA diffusion weights. Changing the planner option without changing the checkpoint is invalid.
Ref2VA supplies only the active original image as persistent, non-positional <Picture 1> reference information. It can preserve identity, texture, clothing, or scene semantics throughout a task, but it does not pin that picture to an exact video frame.
Ref2VA supports:
segment_strategy = warmup_resetreset_anchor = first_onlytechnical_split_source = original_image_reset
It does not support FL2VA hard cuts, first/last bridges, generated-continuation keyframes, or same-first/last endpoints. The planner rejects these contradictory combinations rather than mixing minimax_refs and minimax_keyframes.
ref_image_size controls reference preparation:
matchlimits the reference toward the generation canvas area.maxallows up to a 2048-pixel short edge at greater encoder and sampling cost.
Both preserve aspect ratio and avoid upscaling the source.
Visual Conditioning Strength
Visual Conditioning Strength is one shared root Float control connected to both render subgraphs and accepts values from 0.0 through 1.0.
0.999is the native compatibility setting. Eclipse omits the override so ComfyUI uses H3's implicit native value and original conditioning metadata shape.1.0removes conditioning noise and creates the strongest VisualVAE anchor. It can also increase static behavior or reference-like flashes.Lower values add more seeded noise to the VisualVAE condition and may loosen source adherence. They are not a direct motion control.
This control does not change Qwen visual tokens or audio conditioning. For a useful comparison, keep the audio, images, prompts, seed, frame count, model, LoRA, sampler, scheduler, and steps identical; then compare motion, source similarity, and reference-like flashes.
Choosing the V2 Segment Strategy
Warmup Reset — Recommended Default
Use:
segment_strategy = warmup_reset
warmup_seconds = 2.0
reset_anchor = first_only
At an image transition, the destination starts an independent task with its original image at hidden local frame 0. H3 generates the warmup, the workflow crops the complete warmup, and the first visible destination-owned frame at T is generated rather than pasted from the exact input.
The warmup audio comes from the real preceding timeline range. At the start or end of the complete audio, only the genuinely out-of-range portion is padded.
This mode prevents literal source-frame exposure at ordinary destination changes while giving each image a clean re-anchor.
First/Last Bridge — FL2VA Only
Use:
segment_strategy = first_last_bridge
The source-owned task receives ordered source and destination images and interpolates toward a hidden destination endpoint at T. A separately warmed, destination-owned task supplies the visible generated frame at T.
Use this when a visible transformation or cinematic handoff is more important than keeping the old appearance unchanged until the boundary. Destination influence can develop before T because the pair conditions the bridge task.
This is not the old V1 short-window/full-interval control. V2 has one explicit first/last bridge strategy.
Hard Cut — FL2VA Only
Use:
segment_strategy = hard_cut
The destination's exact source guide is retained at frame T. This produces the strictest image-ownership boundary, but unlike warmup_reset, the literal input frame is intentionally visible at the cut.
Use hard cuts when exact source placement matters more than hiding the reference frame.
Reset Anchor
first_only is the recommended default.
same_first_last_experimental adds a discarded same-source endpoint where the task layout permits it. It never replaces a real bridge destination and is not available in Ref2VA mode.
Long Sections and Technical Seams
What Is a Technical Seam?
H3 render lengths must follow the 17k+5 frame grid and remain between 124 and 362 frames. A long image-owned interval may therefore require multiple tasks. The boundary between those tasks is a technical seam: the image owner does not change, but a new H3 generation begins.
max_render_frames controls the largest allowed task. Lowering it can reduce per-task memory use, but it creates more technical seams and more total sampling work. Hidden warmups and endpoints also consume render capacity.
Plain Original-Image Reset — Neutral Test Preset
Use:
technical_split_source = original_image_reset
technical_seam_style = plain_reset
Each new technical task regenerates independently from the active original image. Its guide and warmup stay hidden, but two independent generations can choose different framing or camera position. A visible same-character framing reset is therefore possible even though no exact input frame was pasted.
Intentional Camera Cut
Use:
technical_split_source = original_image_reset
technical_seam_style = intentional_camera_cut
This keeps the same independent hidden reset, then appends the selected Technical Cut Instructions item only to eligible same-image technical tasks. Connect either a scalar string, String Multiline List [Eclipse] string_list, or Wildcard Processor List [Eclipse] list. Every non-empty list item is one complete direction; blank items are ignored.
Eligible seams take directions in chronological order. When there are more eligible seams than directions, the list cycles from the beginning. For example, three directions are assigned as prompt indices 0, 1, 2, 0, 1, 2, .... A blank or disconnected input uses the built-in direction requesting a clearly different camera angle and shot size while preserving identity, wardrobe, scene, lighting, and ongoing action.
Inline wildcards are another option: place wildcard syntax such as __camera_angle__ inside a scalar or list item and expand it upstream. Inline expansion varies wording within that item; list output controls which complete direction each eligible seam receives.
The instruction is never added to:
The initial task.
A real image transition.
A first/last bridge.
A generated-continuation task.
The source remains at hidden local frame 0, the complete warmup remains cropped, and the first visible frame is generated. The workflow does not rotate, crop, transform, or paste the source image to fake a camera cut.
This option makes a necessary technical discontinuity look editorially deliberate. It remains generative, so edit the instruction when the requested angle or shot size is too weak for your source and prompt.
Generated Continuation
Use:
technical_split_source = generated_continuation
technical_seam_style = plain_reset
The next technical task receives a detached copy of the preceding 22 exact generated frames read from the stored timeline, without loading the full video. This aims for visual continuity across the split, but repeated chaining can accumulate appearance or quality drift over a long timeline.
intentional_camera_cut cannot be combined with generated_continuation; the planner rejects that contradiction.
Understanding H3 Plan Report
Always inspect H3 Plan Report before judging a preview. It reports:
Master-audio duration and exact retained frame count.
Number of selected timeline images.
Even or manual transition targets.
Activity-aligned transition results.
Conditioning family and segment strategy.
Technical split source and seam style.
Every task's visible half-open range
[start, end).Image/prompt ownership, render length, warmup/crop, and guide roles.
The planner also produces analysis_manifest, a JSON string without image or audio tensors. It records transitions, image and technical seams, retained ranges, prompt owners, selected technical-cut prompt indices, guide positions, and expected source ownership. Inspect this manifest to see which cyclic prompt index was assigned to each eligible intentional seam. It is primarily useful for advanced inspection and external analysis; the workflow does not need it to render.
Resolution, Length, and Performance
Choose generation width and height in Smart Folder and keep both divisible by 32. The shared Image Resize node downstream from Image Selector normalizes selected timeline images to that canvas before planning and conditioning, so source files do not need to be manually resized first.
The saved V4 exports at the H3 canvas size. Its timeline is a custom data type, not an IMAGE batch: ordinary IMAGE resize/upscale nodes cannot consume it directly. Avoid rebuilding a full IMAGE batch for a long video, which restores the accumulated RAM cost. For behavior comparisons, keep the generation canvas fixed.
H3 generates legal 17k+5 lengths. The planner chooses the shortest legal render that can supply each planned retained range and balances long regions across the minimum number of tasks.
For lower VRAM use:
Reduce resolution first.
Lower
max_render_framesif one task is still too large.Reduce optional patches or LoRAs.
Test a shorter audio excerpt.
Diff-Aid is a conditioning experiment, not a guaranteed memory or speed optimization. Compare its effect with the same inputs and seed.
Lowering max_render_frames trades memory for more generations and more technical seams. If those seams become visible, choose intentional camera cuts or generated continuation according to the trade-off you prefer.
The disk timeline removes growing retained-frame batches from the loop. Each task still needs memory for sampling and native VAE decoding; it does not add temporal tiling, lower precision, reduce resolution, or shorten planned scenes. Model memory and one task's decode can still exhaust RAM or VRAM.
The accumulated preview streams every completed task and is re-encoded after each extension, so preview time and disk I/O still grow with the timeline. Previews are compressed viewing copies; continuity, caption backgrounds, and clean export read the exact stored pixels, never the preview MP4. The final MP4 uses the chosen encoder settings; exact intermediate storage does not make a lossy H.264 export lossless.
Temporary Disk Space and Cleanup
At 512 × 896, 24 FPS, RGB float32, four minutes of retained frames need about 29.5 GiB on the ComfyUI temporary drive, plus all accumulated previews, caption renders, and export space. Higher resolution or longer audio requires more space. Reserve headroom; a drive with only 30 GiB free is not enough for that four-minute case and its additional files.
By default the temporary folder is ComfyUI/temp; a configured --temp-directory can put it on a different drive. Timeline disk-full errors show the actual folder. Check that drive even if the final output drive has room. Configure a larger temporary drive before starting ComfyUI when needed.
Finishing the queue does not necessarily delete the frames. ComfyUI can keep timeline outputs and loop states cached for caption edits and re-exports. Exact chunks are removed when their final owner is released. Clearing execution caches can release them, but unloading models or completing a save is not a guarantee. Accumulated preview MP4s remain until temporary-file cleanup.
If you need to reclaim space manually:
Keep the final videos you want from
output. Finish or cancel the current job, then stop the ComfyUI backend, not just the browser tab.In its temporary folder, remove unneeded
EclipseFrames_*folders containingframes.npy. These contain the exact cached video pixels and are usually the largest files. Deleting them means later caption edits or re-exports need generation again.Remove unneeded
EclipseVideo_temp_*.mp4accumulated previews. Temporary caption/video copies useEclipseLyrics_*.mp4andEclipsePreview_*.mp4; keep their final exported versions first. An interrupted render may also leaveEclipseLyricsBackground_*temporary excerpt folders.Restart ComfyUI before queueing again. Leave
input,models, and wanted final exports inoutputalone.
Do not delete frame chunks while the backend is running: cached nodes may still reference them. Incomplete chunk writes are removed on failure, but abrupt process termination can leave files. This storage is local to the execution, not restart recovery; leftover chunks and preview MP4s cannot automatically resume a run.
Output
The final stage:
Writes only retained task ranges after removing planned warmups, overlaps, hidden endpoints, and grid padding.
Appends exact chunk references without a crossfade or growing IMAGE batch.
Selects the exact planned video-frame range for final rendering.
Optionally renders lyric captions over those frames with matching excerpt audio.
Saves H.264 MP4 at 24 FPS; the saved export is CRF
19, presetveryfast.Uses the selected master excerpt, preserving the chosen soundtrack rather than replacing it with generated H3 audio or isolated vocals.
Leaves the separately saved full-length generated FLAC unchanged.
Embeds workflow and generation metadata while those save options remain enabled.
Embedded metadata can include prompts, model selections, filenames, and saved local paths. Disable workflow or generation-data embedding before creating a public copy if you do not want that information included.
Troubleshooting
V4 Timeline, Caption, or V2 Nodes Are Missing
Update ComfyUI Eclipse to 4.4.10 or newer, restart ComfyUI, and reload the downloaded workflow. The exact required node names are:
Decode and Append Timeline [Eclipse]Trim Frame Timeline [Eclipse]Frame Timeline Loop Gate [Eclipse]Render Lyric Captions [Eclipse]MiniMax H3 Audio Timeline Planner V2 [Eclipse]MiniMax H3 Audio Plan Step V2 [Eclipse]MiniMax H3 Segmented Conditioning V2 [Eclipse]
Load Audio Has No Incoming Input or Review Stop
Update Eclipse to 4.4.10 or newer and restart ComfyUI. Both source and excerpt nodes must be Load Audio [Eclipse]. Keep the full-song source at start 0, duration 0, and the H3 clip on Incoming audio. For a vanished temporary preview, execute again; do not treat the preview file as the saved song.
The Wrong Song Plays or Both Generators Run
Auto prefers incoming AUDIO. Choose Selected file in Full song source to force a recording, and mute both music generators to stop their independent save nodes from requesting generation. For generated music, enable only the intended branch; YuE2 has priority when both are active. Match the caption lyrics to whichever source was selected.
The Diff-Aid Node Is Missing
Install or update ComfyUI-DiffAid-Patches, restart ComfyUI, and reload the workflow. Muting the group does not remove the custom-node dependency from the saved graph. For a dependency-free baseline, remove the complete DIFF-AID group and keep the final selector connected to MODEL_H3_SOL and MODEL_H3_LORA.
A Model Patch Fails During Sampling
Keep Fused Modulation disabled if its forward function reports an unsupported attention argument. For a Diff-Aid failure involving compiler recording or aimdo, compare with Diff-Aid disabled and, if needed, launch ComfyUI with --disable-comfy-compiler. Recheck compatibility after updating ComfyUI and the patch packages; these workarounds do not guarantee every patch combination.
The Planner Rejects My FL2VA/Ref2VA Settings
FL2VA and Ref2VA are different checkpoint and conditioning families. Ref2VA only accepts warmup_reset, first_only, and original-image technical resets. Return to that combination or switch both the planner and diffusion model back to FL2VA.
The Wrong Image or Prompt Appears
Check the filename sort order, Image Selector thumbnails, and non-empty prompt lines. Blank prompt lines are removed; they do not reserve an image index.
A Real Image Transition Is Too Abrupt
Use
warmup_resetto hide the exact destination frame.Try
first_last_bridgewhen you want an FL2VA interpolation.Use images with more compatible framing and lighting.
Disable activity alignment if an exact manual boundary matters more than a nearby phrase break.
The Exact Source Image Appears at a Transition
Check segment_strategy. This is expected with hard_cut, which intentionally retains the destination guide at T. Use warmup_reset to hide it behind a generated lead-in.
Conditioning Is Static or Reference-Like
Return Visual Conditioning Strength to 0.999 first; that is the native H3 compatibility path. 1.0 is a stronger anchor, not a motion boost, and can increase static behavior or reference-like flashes. Lower values add seeded VisualVAE noise and may loosen adherence rather than directly creating motion.
For a controlled comparison, render the same shortest useful clip with the same seed and all other settings fixed. Compare a lower experimental value, native 0.999, and 1.0; judge motion, source similarity, and flash behavior together. Keep the step count appropriate for the active acceleration LoRA in every candidate.
The Camera Jumps Without an Image Change
That is normally a same-image technical reset:
Keep
plain_resetwhen the reset is acceptable.Choose
intentional_camera_cutto request a deliberate editorial angle/size change.Choose
generated_continuationto prioritize continuity, accepting possible long-run drift.Increase
max_render_frameswhen memory permits to reduce the number of seams.
The Intentional Cut Is Too Weak
Edit Technical Cut Instructions with a concrete angle and shot size appropriate for your current scene. Add more complete list items to alternate directions across technical seams. The feature is prompt-driven, so it requests rather than geometrically forces the camera change.
The Video Timing Does Not Match the Audio
Keep the workflow at 24 FPS. Confirm that all H3 duration/planning consumers and clean-video mux inputs use MASTER_AUDIO_H3 from H3 clip, rather than a generator's full-song output. Caption audio deliberately uses the full song with the shared trim. Master excerpt and the optional H3 vocal stem must begin together and differ in duration by no more than one video frame. For a conditioning sample-rate/count error, check the shared Demucs/preview controls described above.
Audio Demux Is Off but Demucs Still Runs
Confirm Audio Demux controls the Demucs Mode Bridge. Its matching bridge must mute Hybrid Demucs, Preview Audio, and the conditioning-audio Set node. The three fallback selectors must keep conditioning audio first and master audio second; the first active value wins.
The separate Optional full-song vocals node belongs to caption alignment and is not controlled by Audio Demux. Mute it separately to use the original mix for captions, or supply valid corrected timing to skip its lazy analysis input.
The Preview Is Shorter Than the Audio
The base preview contains only task 0. The accumulated preview contains only the tasks completed so far, in order, with the matching audio prefix. Wait for the extension loop and final save. The accumulated preview is intentionally retained in DiskTimeline; it is not reduced to the most recent scene.
Captions Are Missing, Wrong, or Shifted
Confirm H3 Render Lyrics is enabled and the manual-lyrics override is enabled only when it contains the intended full-song text. For file input, do not leave unrelated generator lyrics selected. Check the language and alignment report. Unresolved lines can be omitted; transcription additions can mishear singing. Review and correct the full-song timing JSON as needed.
For a systematic offset, keep Full song source untrimmed, share the same clip start with H3 and captions, and feed H3 clip's actual duration to the renderer. Use a full-song vocal stem for captions. Do not apply the excerpt offset to the already trimmed background a second time.
Caption Model or Font Is Missing
Install Eclipse's requirements in the ComfyUI Python environment and place the five pinned Whisper files in models/whisper/large-v3/, as described above. H3's Qwen and Wav2Vec weights do not replace that model. Select an installed font covering the lyrics' script. For a clean export without alignment, mute the caption renderer; the final selector will use the H3 timeline.
Disk Is Full or Temporary Files Remain After Completion
Follow Temporary Disk Space and Cleanup above. The large EclipseFrames_* directories and accumulated EclipseVideo_temp_*.mp4 previews live on the temporary drive, which can differ from the export drive. Stop the backend before manual deletion. Finishing a run can leave frame chunks cached intentionally for caption changes and re-exports. Deleting them or restarting loses that reuse; there is no automatic restart recovery.
The Workflow Runs Out of Memory
First confirm you loaded the DiskTimeline variant and the loop carries ECLIPSE_FRAME_TIMELINE, not accumulated IMAGE batches. Do not convert the final timeline or VIDEO back to a full IMAGE batch.
For sampling or native decode failures, reduce generation resolution first, then lower max_render_frames, reduce optional model patches or LoRAs, and test a shorter audio segment. More, smaller tasks reduce each task's peak workload but take longer and create more technical seams. A downstream output resize cannot reduce H3's generation-time memory use.
Suggested Timeline Presets
Recommended Hidden-Reset Timeline
conditioning_family = fl2va_keyframes
segment_strategy = warmup_reset
warmup_seconds = 2.0
reset_anchor = first_only
technical_split_source = original_image_reset
technical_seam_style = plain_reset
visual_condition_strength = 0.999
Editorial Technical Cuts
conditioning_family = fl2va_keyframes
segment_strategy = warmup_reset
technical_split_source = original_image_reset
technical_seam_style = intentional_camera_cut
Edit Technical Cut Instructions to suit the subject and current framing. Each non-empty item is a complete direction, and the list cycles across eligible technical seams.
Maximum Technical Continuity
conditioning_family = fl2va_keyframes
segment_strategy = warmup_reset
technical_split_source = generated_continuation
technical_seam_style = plain_reset
Use this when same-image seam continuity matters more than accumulated drift.
FL2VA Image Interpolation
conditioning_family = fl2va_keyframes
segment_strategy = first_last_bridge
technical_split_source = original_image_reset
Use compatible source/destination images and expect destination influence before the exact ownership boundary.
Exact FL2VA Source Cuts
conditioning_family = fl2va_keyframes
segment_strategy = hard_cut
The destination source frame is intentionally visible at T.
Ref2VA Identity/Reference Mode
conditioning_family = ref2va_active_reference
segment_strategy = warmup_reset
reset_anchor = first_only
technical_split_source = original_image_reset
ref_image_size = match
Load the Ref2VA checkpoint before queuing.
Custom Node Packages Used
ComfyUI core supplies H3 latent/guide nodes, schedulers, samplers, native VAE decoding, audio encoders, Music 3 and YuE2 generation, and full-song audio saving. Eclipse supplies the exact frame-timeline transport and caption rendering.
Start with the recommended neutral first-test settings and inspect H3 Plan Report. Once the first transition and technical seam behavior match your intent, increase the audio length or add more timeline images.




