Sign In

MiniMax H3 LongForge FL2VA & REF2VA (LONG VIDEO)

Updated: Sep 27, 2026

base modeleasylatenth3worflowminimax

Download

2 variants available

Type
Workflows
Stats

266

Reviews
Published

Sep 27, 2026

Base Model

MiniMax H3

Hash
AutoV2
250E38BD9E
default creator card background decoration
Uploads - 9

9

Likes - 142

142

Downloads - 4828

4.8K

Generation, training and LoRA distribution on Civitai are covered by Civitai’s own license agreement with MiniMax. If you download these weights and run them yourself, your use is instead governed by the MiniMax H3 Community License Agreement, whose grant excludes the European Union, the United Kingdom, the Republic of Korea and the United States of America.

MiniMax H3

1.png

MiniMax H3 LongForge v3.7 — FL2VA & REF2VA

My Telegram Channels

AI Bob Public
AI DEV LAB

Support the project

If LongForge helps you, you can support future development through the Tip button on my Civitai profile or with an optional donation:

USDT · TRON (TRC20)

THLU3yu8VpoPqP4bMLCp86FRhoCweJ2Qz4

Please send only USDT through the TRON (TRC20) network. Thank you!

One editable ComfyUI workflow for generating, continuing, enhancing, reviewing, recovering and exporting long MiniMax H3 videos with audio.

LongForge transfers protected video and audio latent context between scenes, helping preserve motion, character appearance, camera movement, scene geometry and ambience.

Version 3.7 adds an optional local Prompt Assistant, automatic continuity modes, an AUTO profile for Adaptive Native Cache, and more detailed audio and scene-join diagnostics. HQ Finish, the clean refiner branch, persistent projects, recovery metadata and long-audio references remain part of the same workflow.

Important: Turbo LoRA and Adaptive Native Cache are separate acceleration methods. Keep Native Cache bypassed when using Turbo or another short-step acceleration setup.

What’s new since v3.5

Optional local Prompt Assistant

Write a short scene idea directly in Project & Scenes, then press Generate draft to create an English H3 prompt with a separate local Gemma 4 text model.

The assistant uses:

  • the selected FL2VA / REF2VA pipeline;

  • the configured scene duration;

  • active first/last-frame inputs;

  • available reference labels;

  • the previous scene’s prompt for continuation.

Review or edit the result, then choose Apply to scene. Closing the draft window leaves your scene unchanged. Applying a draft to an already-saved scene changes its editor text; regenerate that scene to update the video.

The bundled base and reference prompt-writing guides provide the required structure. Validation checks required fields, their order, unavailable reference labels and some incorrect reference-use claims.

The assistant receives text and reference labels, not the actual images, video frames or audio. Describe important source features in your idea. In REF2VA, a picture is an appearance/style reference by default; explicitly request a first, last or intermediate keyframe when that is the intended use.

Gemma is optional and separate from the H3 text encoder. Normal video generation and export do not require it.

Automatic continuity modes

Video & Continuity now offers three modes:

ModeBehaviorMANUALUses your selected overlap exactly. The supplied workflow starts at 22 frames.AUTO BALANCEDEvaluates the previous native AV tail and normally chooses 5 / 22 / 39 frames.AUTO SAFEUses the more conservative 22 / 39 / 56-frame pool when enough context and scene length are available.

AUTO examines changes in the saved video latent and, when audio preservation is enabled, the audio latent. It selects a valid overlap for the next scene without rewriting the previous scene. Very short windows can require smaller valid values.

The actual overlap and selection reason are saved with each scene and shown in its diagnostics. Manual controls remain available. Automatic selection helps choose context length; it does not guarantee an invisible or inaudible join.

Adaptive Native Cache — AUTO profile

The new AUTO profile calibrates on the current scene and selects an existing manual profile according to video/audio behavior.

Before allowing its first reuse, and periodically afterward, AUTO computes a fresh model output and checks the proposed cached result against it. Failed checks move the controller toward a more conservative profile; repeated failures can disable reuse for the rest of that scene. Verification steps always keep the fresh result.

The existing QUALITY / VOICE / ACTION / BALANCED / FAST / PREVIEW / DRAFT10 / CUSTOM profiles remain available. The supplied workflow uses BALANCED; AUTO is an optional selection.

Additional compatibility guards detect recognizable acceleration/Turbo LoRAs in the active generation branch and disable reuse for stochastic samplers such as ancestral, SDE, DDPM and LCM variants. A renamed acceleration checkpoint may require manual bypass.

Audio export and join diagnostics

Audio work focuses on identifying the source of a problem and keeping repairs local:

  • Repair audio clicks applies a bounded correction around detected isolated clicks instead of smoothing a large part of the soundtrack.

  • Diagnostics examine both audio channels and distinguish a short transient from a sustained level reset.

  • Export reports identify joins that need review as before Scene N.

  • Diagnostic WAVs · export only saves the continuous decoded audio before seam repair and the final PCM before AAC encoding.

  • FFmpeg export uses the film’s timeline duration to avoid premature termination that could cut off a trailing AAC packet.

Broad dropouts, missing sound and resets already generated by H3 cannot be reconstructed during export. Diagnostic WAVs help determine where a transition appeared; enabling them alone does not remove it.

Updated public workflow and guides

The supplied workflow contains a neutral boat example, an empty Reference Library and placeholders for model selections. Both LoRA loaders and both FL2VA image loaders start bypassed. The KJ attention patches and Native Cache / BALANCED are active, HQ Finish is OFF, and Video metadata is set to NONE.

The English/Russian guide now covers Prompt Assistant setup, sequential scene drafting, automatic continuity, cache AUTO, audio diagnostics and the actual v3.7 starting configuration.

Main features

  • Unified FL2VA and REF2VA workflow.

  • Persistent multi-scene projects.

  • Protected video and audio latent continuation.

  • Manual and automatic overlap selection.

  • Scene Studio with searchable prompts, statuses and previews.

  • Optional local prompt drafting with Gemma 4.

  • Drag-and-drop reordering of pending scenes.

  • Built-in image, video and audio Reference Library.

  • Automatic long-audio slicing for REF2VA.

  • Scene regeneration and controlled variations.

  • Optional recovery metadata inside preview and export files.

  • Integrated native latent upscaling and refinement.

  • Separate clean model branch for refinement.

  • Adaptive Native Cache with manual and AUTO profiles.

  • Local audio-click repair, optional video-seam correction and join diagnostics.

  • Complete film export with generated audio.

  • Per-project locking and atomic updates on Windows.

Generation modes

FL2VA

Supports text-to-video, image-to-video, optional first-frame guidance, optional last-frame guidance and combined first/last-frame generation.

Leave both image loaders bypassed for text-only generation. First frame anchors the film’s beginning. Last frame applies to selects either the last scene currently listed or every scene.

REF2VA

Supports up to:

  • 9 image references;

  • 3 video references with optional paired audio;

  • 3 standalone audio references.

Switch the pipeline inside the same workflow and select the matching MiniMax H3 diffusion model.

FIRST SCENE presents ordinary references only at the beginning; later scenes continue from saved latent context. EVERY SCENE presents them again and can pull the generation back toward the original pose, composition or sound. Follow-scenes audio advances independently of this policy.

Adaptive Native Cache

Adaptive Native Cache reuses selected stable model outputs while keeping the sampler schedule intact. It is intended for native H3 generation, normally 20 or 25 steps with CFG 1.0.

ProfilePurposeAUTOScene-specific calibration, fresh-output checks and automatic fallback.QUALITYThe most conservative manual reuse profile.VOICEStricter audio protection for speech and vocals.ACTIONCautious reuse for motion-heavy scenes.BALANCEDGeneral starting point and supplied workflow default.FASTMore aggressive reuse.PREVIEWFaster drafts and prompt testing.DRAFT10Targets roughly ten reused steps in a 20-step run when checks allow.CUSTOMManual thresholds, budgets, warmup, reuse window and consecutive-reuse limits.

Custom controls appear only when CUSTOM is selected.

Video and audio are evaluated separately; audio can veto reuse. Warmup, the final step and protected cache-window edges use full computation. Reuse is disabled for fewer than 16 steps, CFG other than 1.0 and detected incompatible configurations. Do not stack another output-cache system on the same model branch.

Cache does not change prompt conditioning, saved continuation latents, HQ refinement, preview decoding or film export. Reports show actual full/reused model calls; estimated time savings are not a measured speedup guarantee.

Bypass cache for Turbo, short-step acceleration or an exact native comparison. Compare a chosen profile against bypass with the same seed before relying on it for final output.

Integrated HQ Finish

The optional HQ Finish node processes each completed scene before preview and export.

ModeProcessingOFFNative H3 output.UPSCALELearned MiniMax H3 3D latent upscaling.UPSCALE + REFINELatent upscaling followed by low-strength native H3 video refinement.

The integration is built into LongForge. Download a compatible third-party 3D-conv upscaler checkpoint separately; no separate upscaler node pack is required.

Clean refiner branch

The supplied workflow has two model paths:

  • Generation: shared attention patches → Turbo LoRA → Extra LoRA → optional Adaptive Native Cache.

  • Refinement: shared attention-patched H3 model before both LoRAs and cache.

This keeps generation LoRAs and cache out of the refinement pass while retaining the shared attention patches.

Recommended starting settings:

  • Refine steps: 5

  • Refine denoise: 0.25

  • Sampler: Euler

  • Scheduler: Simple

  • CFG: 1.0

Five passes at 0.25 denoise sample the final part of a native 20-step trajectory. Other useful pairs are 4 / 0.20, 6 / 0.30 and 7 / 0.35.

Native continuation and HQ export

LongForge preserves the native video/audio latent for continuation and stores enhanced HQ video for previews and export. Audio is not upscaled and remains protected during refinement. The previous HQ overlap is carried into the next HQ scene.

Keep the same HQ mode, checkpoint, precision and output geometry throughout the film. Regenerate from Scene 1 when changing these settings.

For example, a film generated at 736 × 416 with a 2.0× HQ scale produces 1472 × 832 video while continuation still uses the original native latent. HQ increases processing time and storage requirements.

Scene Studio

Project & Scenes provides editable scene cards, search, generation statuses, saved previews, pending-scene reordering, regeneration, variation and project cleanup controls.

Pending cards can be dragged into a new order while search is empty. Saved scenes depend on the latent context before them; changing an already-generated part requires rebuilding its continuation.

Open project restores saved prompts and video, sampling and HQ settings. Model and LoRA selections remain in the workflow JSON. Save that JSON before closing the browser to retain pending editor changes and loader selections.

Use a different project name for a separate film. New film resets the active history for the current name while retaining editor text and old takes. Clean unused removes takes outside the active chain; Delete film removes the entire project.

Audio continuity

Recommended starting settings:

  • Continuity mode: MANUAL

  • Overlap: 22 frames

  • Preserve audio overlap: enabled

  • Extra audio history: 0 frames

You can select AUTO BALANCED or AUTO SAFE when you want LongForge to choose overlap from the preceding scene. A 5-frame overlap is supported but provides less context and a higher risk of motion or audio resets.

Extra audio history sits immediately before the protected overlap. If active REF2VA audio references conflict with this additional history, LongForge disables only Extra audio history for that scene and reports it; protected audio overlap remains enabled.

Long-audio references

In REF2VA, choose Audio use → Follow scenes · REF2VA in the audio card. LongForge advances through the source according to the saved film timeline, avoiding a duplicate reference for audio already covered by the protected overlap.

This guides H3’s own audio synthesis. It does not replace the final soundtrack or guarantee exact words, voices, rhythm or timing. Actual source intervals appear in the generation report. If the selected source ends early, the remaining reference window is padded with silence.

Seam repair and diagnostic WAVs

Repair audio clicks targets isolated clicks near joins. A broad dropout or sustained reset needs review and may require regeneration with more protected context.

For an audible join, enable Diagnostic WAVs · export only, select EXPORT FILM and run again. No new diffusion generation is required. The export folder receives:

  • the MP4;

  • .before_repair.wav — continuous decoded H3 audio before seam repair;

  • .before_aac.wav — audio after seam repair and before AAC encoding.

The WAVs retain float32 samples without lossy compression. Both reflect the Fade film edges setting. Compare files with the same export basename around the reported join.

Repair video seams is optional and OFF by default. Fade film edges applies short ramps only at the beginning and end of the film.

Recovery metadata

Set Video metadata → RECOVERY to embed prompts and workflow recovery information in newly created previews and exported MP4s. This can include model/LoRA filenames and reference-media selections.

Drag a recovery MP4 into ComfyUI to reopen its workflow. A scene preview restores that scene’s graph; a full-film export uses the last scene’s model setup with the active film prompts. It does not automatically switch a different LoRA setup for every card.

Recovery does not contain model weights, LoRA files, original reference media or scene latents. Keep the project folder to continue generating after a restart:

ComfyUI/output/longforge_native/<project>/

The public workflow starts with Video metadata → NONE. Select RECOVERY when you want embedded restoration data. NONE affects newly written videos; it does not remove metadata from existing MP4s or delete project data. ComfyUI’s --disable-metadata also disables embedded recovery.

If recovery export fails because of a JSON nan value, select NONE. This does not change generation itself.

Project handling on Windows

LongForge uses per-project locking, atomic JSON updates and protection against simultaneous generation and cleanup. Successfully generated scenes are saved immediately, including during ALL PENDING.

Clean unused and Delete film require confirmation. Avoid editing or generating the same project simultaneously in several browser tabs.

Existing compatible project/take folders can be reused. Install the supplied v3.7 workflow to access the current controls and clean refiner connection.

Workflow structure and defaults

The graph uses visible native model, CLIP, VAE and LoRA loaders; KJNodes attention patches; and LongForge project, continuity, reference, cache, HQ, prompt-assistant and export nodes.

ComponentSupplied v3.7 settingPipelineFL2VAResolution1344 × 768Scene length158 frames at 24 FPSContinuityMANUAL, 22-frame overlapPreserve audio overlapONExtra audio history0Sampling20 steps, res_multistep, simple, CFG 1, denoise 1Turbo LoRA / Extra LoRABypassed; no files selectedFirst / Last frame loadersBypassed; no files selectedKJ attention patchesEnabledAdaptive Native CacheEnabled, BALANCEDHQ FinishOFFPrompt AssistantOptional; text model not connectedReference LibraryEmptyVideo metadataNONEDiagnostic WAVsOFF

Sampling remains editable. Euler/simple is a straightforward starting point for native sampling and cache comparisons; the supplied JSON retains res_multistep as its selected sampler.

Width and height must be multiples of 32. Scene length and overlap follow the H3 grid 5 + 17 × k frames, with overlap shorter than the next scene. There is no fixed 1 MP resolution cap; practical limits depend on the model and hardware.

Scene duration includes overlap. Three 243-frame scenes with 22-frame overlap produce 243 + 221 + 221 = 685 unique frames, approximately 28.54 seconds at 24 FPS.

Download

Download both archives:

  • ComfyUI-H3-LongForge-v3.7-NodePack.zip — custom nodes, web interface, bundled prompt-writing guides and installation README.

  • H3_LongForge-v3.7-Workflow-Prompt-Guide.zip — H3_LongForge_FL2VA_REF2VA.json and the complete PROMPT_GUIDE_RU_EN.md user and prompt guide.

Models, LoRAs, VAEs, Gemma weights and upscaler checkpoints are not included.

LongForge v3.7 requires:

  • a compatible recent ComfyUI build with native MiniMax H3 support and the V3 node API;

  • a matching H3 FL2VA or REF2VA diffusion model;

  • a compatible H3 text encoder;

  • H3 video and audio VAEs;

  • a usable FFmpeg executable or the imageio-ffmpeg fallback.

The supplied graph also uses ComfyUI-KJNodes and compatible SageAttention for its enabled attention patches. These patches are optional to LongForge’s core pipeline; install their dependencies to use the graph as supplied, or remove the patch nodes and reconnect the model path.

Optional features require matching LoRAs, a compatible 3D latent-upscaler checkpoint, or a separate Gemma 4 text model for Prompt Assistant.

MiniMax H3 models

Comfy-Org / MiniMax-H3

Alternative H3 text encoders

INT8 ConvRot · NVFP4

Choose a format compatible with your ComfyUI build and hardware. These encoders serve H3 conditioning; Prompt Assistant uses its own text-generation model.

Optional Prompt Assistant model

Comfy-Org / Gemma 4 text encoders

The bundled guide lists gemma4_12b_int8_convrot.safetensors and gemma4_e4b_it_fp8_scaled.safetensors as compatible options for a ComfyUI build supporting them.

Attention nodes

ComfyUI-KJNodes

Optional latent upscaler

LBH-123-AI / MiniMax H3 Latent Upscaler

Use a compatible minimax_h3_latent_upscaler_3d_conv_v1_*.safetensors checkpoint. This integration supports the 3D version, not the 2D checkpoint.

Reference-media loading and upscaler integration are built into LongForge; neither requires a separate node pack.

Installation — Windows Portable

  1. Close ComfyUI completely.

  2. Remove the previous ComfyUI-H3-LongForge-NodePack folder. Keep your projects under ComfyUI/output/longforge_native/.

  3. Extract the new node pack so this file exists: ComfyUI/custom_nodes/ComfyUI-H3-LongForge-NodePack/__init__.py.

  4. Do not merge versions or leave duplicate LongForge installations in custom_nodes.

  5. Install/update ComfyUI-KJNodes and compatible SageAttention for the supplied attention nodes.

  6. If your installation needs the FFmpeg fallback, run the following from ComfyUI_windows_portable:

python_embeded\python.exe -m pip install -r ComfyUI\custom_nodes\ComfyUI-H3-LongForge-NodePack\requirements.txt
  1. Restart ComfyUI and refresh the browser with Ctrl+F5.

  2. Open H3_LongForge_FL2VA_REF2VA.json from the workflow archive.

  3. Select the H3 model, H3 text encoder, video VAE and audio VAE. Match Pipeline to the model.

  4. Select optional LoRA or image files before enabling their loaders with Ctrl+B. Bypass Native Cache when using Turbo.

Optional Prompt Assistant setup

  1. Use a ComfyUI build with native Gemma 4 and text-generation support.

  2. Download a compatible Gemma encoder into ComfyUI/models/text_encoders/ and restart ComfyUI.

  3. Add a native Load CLIP node in the Prompt Assistant area.

  4. Select the Gemma file and set type = stable_diffusion.

  5. Connect its CLIP output to Prompt Assistant → text_model.

  6. Write a short idea in a scene card, press Generate draft, review the result and choose Apply to scene.

For several connected scenes, apply Scene 1’s draft before drafting Scene 2, then apply Scene 2 before drafting Scene 3. Each request receives the preceding card’s text. Review the actual generated scene separately when deciding how its continuation should develop.

Optional HQ Finish setup

  1. Download a compatible 3D-conv latent-upscaler checkpoint.

  2. Place it in ComfyUI/models/latent_upscale_models/.

  3. Restart ComfyUI and select the checkpoint in HQ Finish.

  4. Choose UPSCALE or UPSCALE + REFINE.

  5. For refinement, retain the supplied clean refiner connection before both LoRAs and cache.

Generate your film

Give the project a unique name. Write one prompt per scene and use + Scene to add the next continuation.

A simple manual prompt structure:

VIDEO: Describe the subject, action, setting and camera.
SOUND: Describe ambience, effects and dialogue.
MUSIC: No music.

These headings are optional. Full native H3 prompts, including the longer structure produced by Prompt Assistant, can also be used.

Continuation prompts should describe the next stage of the same action while preserving the intended subject, setting, camera and sound. Do not put several repeated VIDEO sections in one card and expect separate scenes.

REF2VA reference markers

Use <Picture 1>, <Video 1> and <Audio 1> for available references. Optional aliases can make mixed inputs easier to address:

hero=image1
motion=video1
score=audio1
voice=video1_audio

Then use {hero}, {motion}, {score} or {voice} in the prompt. Define only sources you actually loaded. With FIRST SCENE, later prompts should continue the generated scene without addressing static references that are no longer being supplied.

Main actions

Action Result

GENERATE NEXT + ONE SCENE -> Generates the next pending scene.
GENERATE NEXT + ALL PENDING -> Generates all remaining scenes in order, saving each one.
REGENERATE SELECTED -> Rebuilds from the selected saved scene with its saved seed.
REGENERATE LAST -> Rebuilds the last saved scene with its saved seed.
VARY SELECTED -> Rebuilds the selected scene with its saved seed plus one.
PREVIEW SELECTED -> Recreates a missing preview from saved latents without diffusion.
EXPORT FILM -> Assembles the active saved scenes into one MP4 with generated audio.

When regenerating, ONE SCENE replaces only the chosen scene and leaves its continuation pending. ALL PENDING also rebuilds the following scenes.

If all cards are already saved, GENERATE NEXT regenerates the selected scene. Add a new card first when you want to extend the film. Export uses saved scenes and does not generate pending cards.

Optional final post-processing

For separate enhancement or frame interpolation after export:

DLSS 5 Visual Enhancer · Downloads

This is a separate application with its own requirements and is not required by LongForge.