Sign In

MiniMax Music Video in ComfyUI | Build full pipeline with MiniMax H3 and Music 3 - v1.0 Showcase

Loading Images

MiniMax H3 and MiniMax Music 3 pipeline turns any song into synced cinematic scenes.

Who it's for: creators who want this pipeline in ComfyUI without assembling nodes from scratch. Not for: one-click results with zero tuning - you still choose inputs, prompts, and settings.

Open preloaded workflow on RunComfy

Open preloaded workflow on RunComfy (browser)

Why RunComfy first
- Fewer missing-node surprises - run the graph in a managed environment before you mirror it locally.
- Quick GPU tryout - useful if your local VRAM or install time is the bottleneck.
- Matches the published JSON - the zip follows the same runnable workflow you can open on RunComfy.

When downloading for local ComfyUI makes sense - you want full control over models on disk, batch scripting, or offline runs.

How to use (local ComfyUI)
1. Load inputs (images/video/audio) in the marked loader nodes.
2. Set prompts, resolution, and seeds; start with a short test run.
3. Export from the Save / Write nodes shown in the graph.

Expectations - First run may pull large weights; cloud runs may require a free RunComfy account.


Overview

Turn your full song into a connected AI visual sequence. Use MiniMax H3 and MiniMax Music 3 to match scenes to rhythm, vocals, and mood. Guide shots with prompts, images, audio, or video references. Keep characters and style consistent. Build performance or narrative clips faster. Export a cohesive cinematic result.

Important nodes:

Key nodes in Comfyui MiniMax music video workflow

MiniMaxH3ReferenceToVideo (#136)

Creates the conditioning that ties prompt, references, and target geometry to a latent for sampling. Use it to steer identity and art direction with a small set of high‑quality images and clear shot notes. If a shot needs stronger adherence to a face or outfit, keep the references stylistically similar and avoid extreme crops. Adjust prompt phrasing before changing sampling settings, as this node most strongly influences who and what appears on screen.

MiniMaxH3SigmaShift (#5960)

Offsets sigma schedules in the MiniMax H3 stack to balance temporal change and audio alignment. Increase the video shift when you want more motion and visual evolution inside a shot; increase the audio shift to favor tight rhythm and lip feel. Use modest changes and test on a short clip before committing an entire scene.

SamplerCustomAdvanced (#125)

Runs the denoising loop with the selected scheduler and guider. This is where you trade time for quality and texture. For rapid ideation, reduce steps and keep seeds fixed; for finals, give it more steps and nudge guidance to refine detail without losing identity. When results overshoot your brief, first simplify the prompt or reduce conflicting references.

VHS_VideoCombine (#5645)

Assembles decoded frames with the chosen audio into a single video file. Keep frame rate constant across all clips meant to be edited together. If visuals drift against the beat, confirm the shot length and fps are consistent with the source track. Trim in your NLE only after renders match tempo and lyric cues inside the workflow.

MiniMaxMusic3TextEncode (#6157)

Encodes the caption and lyrics that define structure, arrangement, and vocal character for MiniMax Music 3. Write captions in sections that cover global metadata, vocal details, and arrangement so the model understands intent. Use lyric section tags like [Intro], [Verse], and [Chorus] to mark transitions the generator can follow. For prompt help, see the official repository’s guidance and tools. MiniMax‑AI/MiniMax‑Music3

Notes

MiniMax Music Video in ComfyUI | Build full pipeline with MiniMax H3 and Music 3 - see RunComfy page for the latest node requirements.

Comments