Sign In

MiniMax H3 - LipSync

Download

1 variant available

Config Other

MiniMaxH3_Audio_Planned_Timeline_N2.json

239.33 KB

Verified:

Type
Workflows
Stats

239

Reviews
Published

Sep 13, 2026

Base Model

MiniMax H3

Hash
AutoV2
C2F1A775A2
default creator card background decoration
Followers - 1504

1.5K

Likes - 174

174

Downloads - 3833

3.8K

Comfy Workflow Badge

MiniMax H3 is licensed by MiniMax under the MiniMax H3 Community License Agreement. That agreement’s Applicable Territory excludes the European Union, the United Kingdom, the Republic of Korea and the United States of America. Your use of H3 and of any H3 derivative is subject to that agreement and its Acceptable Use Policy.

MiniMax H3

MiniMax H3 LipSync V2.0.2 — User Guide

Long-form image-to-video driven by your own audio · one or more timeline images · FL2VA or Ref2VA conditioning · optional vocal-stem guidance · silence-aware transitions · rotating technical-cut prompts · shared visual-conditioning strength


What Is This Workflow?

This workflow turns an external audio track and one or more images into a continuous MiniMax H3 video. The audio determines the exact duration. H3 renders the performance in manageable tasks, the workflow keeps only each task's planned frames, joins them directly, trims the final timeline, and muxes the untouched master audio once at the end.

Use it for:

  • A single character singing, speaking, or performing through a long track.

  • Multiple images that take ownership of different parts of the timeline.

  • Evenly spaced or manually timed image changes.

  • Transitions aligned to pauses and resumed speech or singing.

  • Independent source-image resets, FL2VA first/last interpolation, or exact hard cuts.

  • Long same-image sections that either continue visually, reset plainly, or turn a necessary generation split into an intentional camera cut.

  • FL2VA positional keyframes or Ref2VA semantic identity/reference conditioning.

The base preview shows the first completed task. The accumulated preview grows after every extension, so you can inspect the result before the full timeline is finished.

Recommended first test: use a short audio excerpt, one image, FL2VA, warmup_reset, the default two-second warmup, and native 0.999 visual conditioning strength.


What Changed in V2?

This guide covers segmented workflow revision 2.0.2.

V2 uses three dedicated Eclipse nodes:

  • MiniMax H3 Audio Timeline Planner V2 builds the complete frame-accurate task plan and a tensor-free analysis manifest.

  • MiniMax H3 Audio Plan Step V2 resolves each task's audio, original image, optional generated continuity, guides, legal render length, crop, and retained frame count.

  • MiniMax H3 Segmented Conditioning V2 creates either FL2VA positional keyframes or one persistent Ref2VA active-image reference. It never combines both conditioning families.

The older transition_mode, bridge_span, bridge_window_seconds, and cropped_lookahead_conditioning controls do not belong to V2. Their replacements are segment_strategy, warmup_seconds, conditioning_family, technical_split_source, and technical_seam_style.

V2 also makes hidden warmups, task ownership, prompt ownership, guide positions, technical splits, and expected image ownership explicit. No crossfade conceals a bad seam.

Workflow revision 2.0.2 moves manual transition times and technical-cut directions out of the planner's saved widgets and into visible, connected text nodes. It also adds one shared Visual Conditioning Strength control for both the base and extension render paths. Output resizing now happens downstream of the source-image controls, so changing the output size no longer reloads the folder images or Image Selector. Eclipse automatically converts the old saved widget values when an earlier workflow is opened; it does not replace an input that is already connected.


Before You Run It

Show the Native 0.999 Value Correctly

To enter and see 0.999 exactly, open ComfyUI Settings, go to LiteGraph → Node Widget, enable Disable default float widget rounding., and set Float widget rounding decimal places [0 = auto]. to 3. Reload the page after changing either setting, as required by ComfyUI. These settings only control Float-widget entry and display precision; they do not add noise to the image or change H3 conditioning by themselves.

Update ComfyUI and Required Custom Nodes

MiniMax H3 support must be available in your ComfyUI installation. If the core H3 nodes are missing, update ComfyUI before troubleshooting the workflow.

Install or update these custom-node packages:

Restart ComfyUI after updating, then load the downloaded workflow again.

If any V2 planner, Plan Step, or Segmented Conditioning node is missing, update Eclipse and restart ComfyUI. Loading an older saved workflow again is not enough if the backend has not been updated.

Required Models

Place the files in their matching ComfyUI model folders.

Default FL2VA diffusion model — ComfyUI/models/diffusion_models/

Optional Ref2VA diffusion model — ComfyUI/models/diffusion_models/

Text encoder — ComfyUI/models/text_encoders/

VAEs — ComfyUI/models/vae/

Activity encoder — ComfyUI/models/audio_encoders/

Choose the encoder that matches the spoken language:

Select the activity encoder that matches the spoken language. This encoder only finds quiet gaps and resumed activity; it never replaces either audio stream.

Match Sampling Steps to the Acceleration LoRA

The LoRA stack is optional. If you enable an H3 acceleration LoRA, use the step count and other sampler settings recommended for that exact LoRA. Different acceleration LoRAs target different schedules, so a step count copied from another LoRA can reduce motion, prompt response, or image fidelity.

If you disable acceleration, switch to sampling settings appropriate for the base checkpoint. Keep the selected LoRA and step count fixed when comparing conditioning or transition controls.

Select Your Inputs

Before queuing, select your own master audio and timeline images, verify their order, and enter one visual prompt or one prompt per image. Then confirm that the checkpoint family, text encoder, VAEs, optional LoRAs, sampler, scheduler, step count, output size, and output location match your installation.

Keep the workflow at 24 FPS, and keep generation width and height divisible by 32. Recheck all file and folder selectors before sharing a workflow or embedded metadata.


Quick Start

  1. Load a short audio excerpt.

  2. Leave Audio Demux enabled for vocal-stem conditioning, or disable it to use the master mix directly.

  3. Select one or more timeline images and confirm their order.

  4. Enter one visual prompt, or one non-empty prompt per image.

  5. For the recommended neutral first test, use:

    • conditioning_family = fl2va_keyframes

    • segment_strategy = warmup_reset

    • warmup_seconds = 2.0

    • reset_anchor = first_only

    • technical_split_source = original_image_reset

    • technical_seam_style = plain_reset

    • max_render_frames = 362

  6. Leave Manual Transition Times disconnected or blank for even ownership, or connect one scalar string with one comma-separated time per image after the first.

  7. Keep Visual Conditioning Strength at 0.999 for native H3 compatibility.

  8. Match the sampling step count to the selected acceleration LoRA, or use base-checkpoint settings when no acceleration LoRA is active.

  9. Enable activity-gap alignment if transitions should move toward nearby quiet gaps and resumed activity.

  10. Read H3 Plan Report, then queue the workflow.

The final file is saved only after the extension loop finishes and the exact planned frame count has been retained.

A long timeline can require many H3 generations. Test image order, prompts, checkpoint family, and transition behavior on a short excerpt before committing to a full song or dialogue track.

Model Patch Groups and Fallbacks

The workflow keeps optional model modifications in separate groups. The final selector prefers the most downstream active branch and falls back through the patch chain to the loader/LoRA model when groups are muted.

Scheduled attention, fused modulation, feed-forward chunking, and Diff-Aid are independent experiments, not universal quality presets. Compare optional patches with identical inputs and seed. To establish a clean baseline, mute the complete optional group rather than bypassing one node while leaving its branch selected.


Audio and Prompt Setup

Master Audio

Load Master Audio determines the exact video duration and supplies the final soundtrack. H3-generated audio is discarded. The original master is muxed after the frame timeline is complete.

Describe the visible performance rather than asking H3 to compose a soundtrack:

A natural cinematic performance. The subject sings in time with the supplied
external audio, with expressive but realistic facial movement, stable identity,
coherent body motion, soft stage lighting, and controlled camera movement.

For instrumental passages, consider adding:

Keep relaxed closed lips during instrumental sections and sing only when vocals
are audible.

Visual prompting and model behavior can still produce occasional mouth motion; vocal separation is not a guaranteed lip lock.

One Prompt or One Prompt Per Image

Prompt Per Image (Blank Lines Ignored) maps non-empty lines to timeline images:

  • One non-empty line applies everywhere through prompt-index fallback.

  • With multiple lines, line 1 belongs to image 1, line 2 to image 2, and so on.

  • Blank lines are ignored and do not reserve an image position.

  • If an image has no matching line, the workflow falls back to the first prompt.

Each task receives one complete prompt. A first/last bridge from image A to image B keeps prompt A; prompt B begins with B-owned output at transition frame T. Prompts never change halfway through one generation task.

Optional Demucs Vocal Conditioning

Audio Demux controls the Hybrid Demucs branch.

When enabled:

  • The untouched master enters Hybrid Demucs.

  • The raw vocal stem conditions both H3 sampling paths.

  • The same vocal stem feeds Wav2Vec activity analysis.

  • Preview Audio lets you hear the extracted stem.

When disabled, the workflow's three fallback selectors use the master audio for:

  1. The base H3 task.

  2. Every extension H3 task.

  3. Wav2Vec activity analysis.

The master always controls duration and the final mux. A separate conditioning stem must begin at the same time and may differ in duration by no more than one 24 FPS frame. Conditioning slices keep their native sample rate and are padded only where a task boundary requires it.


Timeline Images and Transition Timing

Image Order

The selected image batch is the timeline order:

  • One image remains the active identity/reference for the complete video.

  • Two or more images divide the timeline into ownership intervals.

  • Transition frame T belongs to the destination image.

Review the thumbnails carefully. Source-folder filename sorting can change the order you expected.

Compatible identities, composition, lighting, and aspect ratios are easier to connect. Large differences are allowed, but they ask H3 to invent a more difficult handoff.

Manual Transition Times

Enter one comma-separated time in seconds for every image after the first in the connected String Multiline [Eclipse] node. Keep this as one unchanged string; the multiline node makes it easy to inspect and edit, but it does not change the comma-separated parser.

For four images:

8, 17.5, 26

This assigns image 2 near 8 seconds, image 3 near 17.5 seconds, and image 4 near 26 seconds. Values must be strictly increasing and remain inside the audio.

Leave the text node blank, or disconnect the planner input, to divide the complete audio evenly among the selected images.

When activity alignment is enabled, even or manual times are search targets. The planner may move them to a nearby qualifying gap. Disable alignment when manual times must remain exact.

Activity-Gap Alignment

The recommended activity-alignment settings are:

  • align_to_activity_gap = enabled

  • transition_edge = activity_resume

  • search_window_seconds = 5.0

  • min_gap_duration = 0.25

  • resume_hold_duration = 0.15

Enable alignment when timing targets may move to nearby qualifying gaps. Leave it disabled when manual transition times must remain exact.

activity_resume places the ownership change when sustained voice activity returns after a pause. silence_start places it when the pause begins.

For multiple images, search_window_seconds controls how far before and after each target the planner searches. If no qualifying gap is found, it keeps the original target.

With one image, there are no image transitions. The planner may instead use eligible quiet gaps as same-image technical split points. If none is available before a task reaches its allowed size, it splits at the maximum legal point.


Choosing the V2 Conditioning Family

FL2VA Keyframes — Default

Use:

conditioning_family = fl2va_keyframes

FL2VA gives H3 positional start and optional end images. The active original is tokenized as <Picture 1>. A real bridge tokenizes source and destination in that order.

Use FL2VA when you need:

  • A strong opening-image anchor.

  • Exact destination images at hard cuts.

  • First/last interpolation between two timeline images.

  • Optional 22-frame generated continuation at technical splits.

Keep the Smart Model Loader on matching FL2VA diffusion weights.

Ref2VA Active Reference

Use:

conditioning_family = ref2va_active_reference

Then load matching Ref2VA diffusion weights. Changing the planner option without changing the checkpoint is invalid.

Ref2VA supplies only the active original image as persistent, non-positional <Picture 1> reference information. It can preserve identity, texture, clothing, or scene semantics throughout a task, but it does not pin that picture to an exact video frame.

Ref2VA supports:

  • segment_strategy = warmup_reset

  • reset_anchor = first_only

  • technical_split_source = original_image_reset

It does not support FL2VA hard cuts, first/last bridges, generated-continuation keyframes, or same-first/last endpoints. The planner rejects these contradictory combinations rather than mixing minimax_refs and minimax_keyframes.

ref_image_size controls reference preparation:

  • match limits the reference toward the generation canvas area.

  • max allows up to a 2048-pixel short edge at greater encoder and sampling cost.

Both preserve aspect ratio and avoid upscaling the source.

Visual Conditioning Strength

Visual Conditioning Strength is one shared root Float control connected to both render subgraphs and accepts values from 0.0 through 1.0.

  • 0.999 is the native compatibility setting. Eclipse omits the override so ComfyUI uses H3's implicit native value and original conditioning metadata shape.

  • 1.0 removes conditioning noise and creates the strongest VisualVAE anchor. It can also increase static behavior or reference-like flashes.

  • Lower values add more seeded noise to the VisualVAE condition and may loosen source adherence. They are not a direct motion control.

This control does not change Qwen visual tokens or audio conditioning. For a useful comparison, keep the audio, images, prompts, seed, frame count, model, LoRA, sampler, scheduler, and steps identical; then compare motion, source similarity, and reference-like flashes.


Choosing the V2 Segment Strategy

Use:

segment_strategy = warmup_reset
warmup_seconds = 2.0
reset_anchor = first_only

At an image transition, the destination starts an independent task with its original image at hidden local frame 0. H3 generates the warmup, the workflow crops the complete warmup, and the first visible destination-owned frame at T is generated rather than pasted from the exact input.

The warmup audio comes from the real preceding timeline range. At the start or end of the complete audio, only the genuinely out-of-range portion is padded.

This mode prevents literal source-frame exposure at ordinary destination changes while giving each image a clean re-anchor.

First/Last Bridge — FL2VA Only

Use:

segment_strategy = first_last_bridge

The source-owned task receives ordered source and destination images and interpolates toward a hidden destination endpoint at T. A separately warmed, destination-owned task supplies the visible generated frame at T.

Use this when a visible transformation or cinematic handoff is more important than keeping the old appearance unchanged until the boundary. Destination influence can develop before T because the pair conditions the bridge task.

This is not the old V1 short-window/full-interval control. V2 has one explicit first/last bridge strategy.

Hard Cut — FL2VA Only

Use:

segment_strategy = hard_cut

The destination's exact source guide is retained at frame T. This produces the strictest image-ownership boundary, but unlike warmup_reset, the literal input frame is intentionally visible at the cut.

Use hard cuts when exact source placement matters more than hiding the reference frame.

Reset Anchor

first_only is the recommended default.

same_first_last_experimental adds a discarded same-source endpoint where the task layout permits it. It never replaces a real bridge destination and is not available in Ref2VA mode.


Long Sections and Technical Seams

What Is a Technical Seam?

H3 render lengths must follow the 17k+5 frame grid and remain between 124 and 362 frames. A long image-owned interval may therefore require multiple tasks. The boundary between those tasks is a technical seam: the image owner does not change, but a new H3 generation begins.

max_render_frames controls the largest allowed task. Lowering it can reduce per-task memory use, but it creates more technical seams and more total sampling work. Hidden warmups and endpoints also consume render capacity.

Plain Original-Image Reset — Default

Use:

technical_split_source = original_image_reset
technical_seam_style = plain_reset

Each new technical task regenerates independently from the active original image. Its guide and warmup stay hidden, but two independent generations can choose different framing or camera position. A visible same-character framing reset is therefore possible even though no exact input frame was pasted.

Intentional Camera Cut

Use:

technical_split_source = original_image_reset
technical_seam_style = intentional_camera_cut

This keeps the same independent hidden reset, then appends the selected Technical Cut Instructions item only to eligible same-image technical tasks. Connect either a scalar string, String Multiline List [Eclipse] string_list, or Wildcard Processor List [Eclipse] list. Every non-empty list item is one complete direction; blank items are ignored.

Eligible seams take directions in chronological order. When there are more eligible seams than directions, the list cycles from the beginning. For example, three directions are assigned as prompt indices 0, 1, 2, 0, 1, 2, .... A blank or disconnected input uses the built-in direction requesting a clearly different camera angle and shot size while preserving identity, wardrobe, scene, lighting, and ongoing action.

Inline wildcards are another option: place wildcard syntax such as __camera_angle__ inside a scalar or list item and expand it upstream. Inline expansion varies wording within that item; list output controls which complete direction each eligible seam receives.

The instruction is never added to:

  • The initial task.

  • A real image transition.

  • A first/last bridge.

  • A generated-continuation task.

The source remains at hidden local frame 0, the complete warmup remains cropped, and the first visible frame is generated. The workflow does not rotate, crop, transform, or paste the source image to fake a camera cut.

This option makes a necessary technical discontinuity look editorially deliberate. It remains generative, so edit the instruction when the requested angle or shot size is too weak for your source and prompt.

Generated Continuation

Use:

technical_split_source = generated_continuation
technical_seam_style = plain_reset

The next technical task receives the preceding 22 generated frames. This aims for visual continuity across the split, but repeated chaining can accumulate appearance or quality drift over a long timeline.

intentional_camera_cut cannot be combined with generated_continuation; the planner rejects that contradiction.


Understanding H3 Plan Report

Always inspect H3 Plan Report before judging a preview. It reports:

  • Master-audio duration and exact retained frame count.

  • Number of selected timeline images.

  • Even or manual transition targets.

  • Activity-aligned transition results.

  • Conditioning family and segment strategy.

  • Technical split source and seam style.

  • Every task's visible half-open range [start, end).

  • Image/prompt ownership, render length, warmup/crop, and guide roles.

The planner also produces analysis_manifest, a JSON string without image or audio tensors. It records transitions, image and technical seams, retained ranges, prompt owners, selected technical-cut prompt indices, guide positions, and expected source ownership. Inspect this manifest to see which cyclic prompt index was assigned to each eligible intentional seam. It is primarily useful for advanced inspection and external analysis; the workflow does not need it to render.


Resolution, Length, and Performance

Choose generation width and height in Smart Folder and keep both divisible by 32. The shared Image Resize node downstream from Image Selector normalizes selected timeline images to that canvas before planning and conditioning, so source files do not need to be manually resized first.

A resize or upscale added after decoded tasks are stitched changes only the delivered output pixels. It does not change H3 conditioning strength, transition ownership, or generated motion. For behavior comparisons, keep the generation canvas fixed and apply the same downstream output resize to every candidate.

H3 generates legal 17k+5 lengths. The planner chooses the shortest legal render that can supply each planned retained range and balances long regions across the minimum number of tasks.

For lower VRAM use:

  1. Reduce resolution first.

  2. Lower max_render_frames if one task is still too large.

  3. Reduce optional patches or LoRAs.

  4. Test a shorter audio excerpt.

Diff-Aid does not lower VRAM use or sampling time. Mute it only to compare its conditioning effect or to remove the optional patch, not as a memory strategy.

Lowering max_render_frames trades memory for more generations and more technical seams. If those seams become visible, choose intentional camera cuts or generated continuation according to the trade-off you prefer.

The accumulated preview is re-encoded after each extension, so preview overhead also grows with the timeline. Downstream output resizing cannot reduce generation-time memory use.


Output

The final stage:

  • Removes planned warmups, overlaps, hidden endpoints, and grid padding.

  • Appends retained task ranges directly without a crossfade.

  • Trims to the exact planned video-frame count.

  • Saves H.264 MP4 at 24 FPS.

  • Muxes the untouched master audio once.

  • Embeds workflow and generation metadata while those save options remain enabled.

Embedded metadata can include prompts, model selections, filenames, and saved local paths. Disable workflow or generation-data embedding before creating a public copy if you do not want that information included.


Troubleshooting

V2 Nodes Are Missing

Update ComfyUI Eclipse to 4.3.44 or newer, restart ComfyUI, and reload the downloaded workflow. The exact required node names are:

  • MiniMax H3 Audio Timeline Planner V2 [Eclipse]

  • MiniMax H3 Audio Plan Step V2 [Eclipse]

  • MiniMax H3 Segmented Conditioning V2 [Eclipse]

The Diff-Aid Node Is Missing

Install or update ComfyUI-DiffAid-Patches, restart ComfyUI, and reload the workflow. Muting the group does not remove the custom-node dependency from the saved graph. For a dependency-free baseline, remove the complete DIFF-AID group and keep the final selector connected to MODEL_H3_SOL and MODEL_H3_LORA.

The Planner Rejects My FL2VA/Ref2VA Settings

FL2VA and Ref2VA are different checkpoint and conditioning families. Ref2VA only accepts warmup_reset, first_only, and original-image technical resets. Return to that combination or switch both the planner and diffusion model back to FL2VA.

The Wrong Image or Prompt Appears

Check the filename sort order, Image Selector thumbnails, and non-empty prompt lines. Blank prompt lines are removed; they do not reserve an image index.

A Real Image Transition Is Too Abrupt

  • Use warmup_reset to hide the exact destination frame.

  • Try first_last_bridge when you want an FL2VA interpolation.

  • Use images with more compatible framing and lighting.

  • Disable activity alignment if an exact manual boundary matters more than a nearby phrase break.

The Exact Source Image Appears at a Transition

Check segment_strategy. This is expected with hard_cut, which intentionally retains the destination guide at T. Use warmup_reset to hide it behind a generated lead-in.

Conditioning Is Static or Reference-Like

Return Visual Conditioning Strength to 0.999 first; that is the native H3 compatibility path. 1.0 is a stronger anchor, not a motion boost, and can increase static behavior or reference-like flashes. Lower values add seeded VisualVAE noise and may loosen adherence rather than directly creating motion.

For a controlled comparison, render the same shortest useful clip with the same seed and all other settings fixed. Compare a lower experimental value, native 0.999, and 1.0; judge motion, source similarity, and flash behavior together. Keep the step count appropriate for the active acceleration LoRA in every candidate.

The Camera Jumps Without an Image Change

That is normally a same-image technical reset:

  • Keep plain_reset when the reset is acceptable.

  • Choose intentional_camera_cut to request a deliberate editorial angle/size change.

  • Choose generated_continuation to prioritize continuity, accepting possible long-run drift.

  • Increase max_render_frames when memory permits to reduce the number of seams.

The Intentional Cut Is Too Weak

Edit Technical Cut Instructions with a concrete angle and shot size appropriate for your current scene. Add more complete list items to alternate directions across technical seams. The feature is prompt-driven, so it requests rather than geometrically forces the camera change.

The Video Timing Does Not Match the Audio

Keep the workflow at 24 FPS. Confirm the master audio and optional vocal stem begin together and differ in duration by no more than one video frame.

Audio Demux Is Off but Demucs Still Runs

Confirm Audio Demux controls the Demucs Mode Bridge. Its matching bridge must mute Hybrid Demucs, Preview Audio, and the conditioning-audio Set node. The three fallback selectors must keep conditioning audio first and master audio second; the first active value wins.

The Preview Is Shorter Than the Audio

The base preview contains only task 0. The accumulated preview contains only the tasks completed so far. Wait for the extension loop and final save.

The Workflow Runs Out of Memory

Reduce generation resolution first, then lower max_render_frames, reduce optional model patches or LoRAs, and test a shorter audio segment. More, smaller tasks reduce each task's peak workload but take longer and create more technical seams. A downstream output resize cannot reduce H3's generation-time memory use.


Suggested V2 Presets

conditioning_family = fl2va_keyframes
segment_strategy = warmup_reset
warmup_seconds = 2.0
reset_anchor = first_only
technical_split_source = original_image_reset
technical_seam_style = plain_reset
visual_condition_strength = 0.999

Editorial Technical Cuts

conditioning_family = fl2va_keyframes
segment_strategy = warmup_reset
technical_split_source = original_image_reset
technical_seam_style = intentional_camera_cut

Edit Technical Cut Instructions to suit the subject and current framing. Each non-empty item is a complete direction, and the list cycles across eligible technical seams.

Maximum Technical Continuity

conditioning_family = fl2va_keyframes
segment_strategy = warmup_reset
technical_split_source = generated_continuation
technical_seam_style = plain_reset

Use this when same-image seam continuity matters more than accumulated drift.

FL2VA Image Interpolation

conditioning_family = fl2va_keyframes
segment_strategy = first_last_bridge
technical_split_source = original_image_reset

Use compatible source/destination images and expect destination influence before the exact ownership boundary.

Exact FL2VA Source Cuts

conditioning_family = fl2va_keyframes
segment_strategy = hard_cut

The destination source frame is intentionally visible at T.

Ref2VA Identity/Reference Mode

conditioning_family = ref2va_active_reference
segment_strategy = warmup_reset
reset_anchor = first_only
technical_split_source = original_image_reset
ref_image_size = match

Load the Ref2VA checkpoint before queuing.


Custom Node Packages Used

ComfyUI core supplies the H3 latent, guide, scheduler, sampler, VAE decode, audio encoder, and standard image-batch functionality.


Start with the recommended neutral first-test settings and inspect H3 Plan Report. Once the first transition and technical seam behavior match your intent, increase the audio length or add more timeline images.