Sign In

MiniMax H3 LongForge FL2VA & REF2VA

Download

2 variants available

Type
Workflows
Stats

89

Reviews
Published

Sep 20, 2026

Base Model

MiniMax H3

Hash
AutoV2
5C29ECA6CC
default creator card background decoration
Uploads - 7

7

Likes - 104

104

Downloads - 3521

3.5K

MiniMax H3 is licensed by MiniMax under the MiniMax H3 Community License Agreement. That agreement’s Applicable Territory excludes the European Union, the United Kingdom, the Republic of Korea and the United States of America. Your use of H3 and of any H3 derivative is subject to that agreement and its Acceptable Use Policy.

MiniMax H3

MiniMax H3 LongForge v3.0 — FL2VA & REF2VA

Support the project

If LongForge helps you, you can support future development through the Tip button on my Civitai profile or an optional donation:

USDT · TRON (TRC20)
THLU3yu8VpoPqP4bMLCp86FRhoCweJ2Qz4

Please send only USDT through the TRON (TRC20) network. Thank you!


One editable ComfyUI workflow to generate, continue, enhance, review, recover and export long MiniMax H3 videos with synchronized audio.

LongForge transfers protected video and audio latent context between scenes to preserve motion, appearance, camera movement, scene geometry and ambience.

Version 3.0 is a major upgrade over v2.0. It adds integrated latent upscaling, optional clean-model refinement, workflow recovery from MP4 files, scene reordering, improved audio continuity, safer project storage and multiple Windows reliability fixes.

What’s new in v3.0

Integrated HQ Finish

A new optional HQ Finish node processes every completed scene before preview and export:

  • OFF — native H3 generation only.

  • UPSCALE — learned MiniMax H3 3D latent upscaling.

  • UPSCALE + REFINE — latent upscaling followed by low-strength native H3 refinement.

The external upscaler node pack is not required. LongForge contains the necessary integration, but the compatible third-party upscaler checkpoint must be downloaded separately.

Clean refiner branch without LoRAs

The supplied workflow now has two clearly separated model branches:

  • Generation branch: shared attention patches → Turbo LoRA → Extra LoRA.

  • Refiner branch: shared attention patches → clean H3 model before both LoRAs.

This means Turbo and other generation LoRAs do not affect the refinement pass. Both branches can still use the same KJNodes Sage Attention patches.

The recommended starting point is:

  • Refine passes: 5

  • Refine strength: 0.25

  • Sampler: Euler

  • Scheduler: Simple

  • CFG: 1.0

Five refine passes at strength 0.25 represent five actual samples over the final part of a 20-step native trajectory. This is not another full 20-step generation.

Other useful matched combinations are:

  • 4 passes / 0.20

  • 6 passes / 0.30

  • 7 passes / 0.35

Native continuation with HQ delivery

LongForge keeps two separate scene representations:

  • the untouched native video/audio latent used to continue the next scene;

  • the enhanced HQ video used for previews and final export.

The refiner never becomes the continuation source. This prevents upscale or refine changes from accumulating across the film and damaging character identity, motion or scene geometry.

Audio never enters the latent upscaler and is protected during refinement. If every scene uses the same HQ configuration, the final exported film is assembled from the upscaled/refined versions.

The protected high-resolution overlap from the previous scene is restored in the next HQ scene, allowing the final export to cut the same timeline without adding a second copy of the overlap.

Scene reordering

Pending scene cards can now be rearranged by dragging them in Scene Studio.

Prompts and scene names move together, so a new scene can be inserted without manually copying and rewriting every following prompt.

Saved scenes cannot be freely shuffled because each one depends on the latent context of the scenes before it. To change an already-generated part of the film, shorten the active continuation at that point, reorder or add cards, and regenerate the affected scenes.

Recovery metadata

LongForge can store a recovery snapshot inside newly created previews and exported MP4 files.

With Video metadata → RECOVERY, the video can contain:

  • the ComfyUI graph;

  • scene prompts;

  • workflow controls;

  • model and LoRA filenames;

  • reference-media selections.

Drag a recovery MP4 back into ComfyUI to reopen its workflow.

Recovery metadata does not contain model files, source media or saved scene latents. Keep the corresponding ComfyUI/output/longforge_native/<project>/ folder if you want to continue generating the film after restarting ComfyUI.

Select Video metadata → NONE before creating videos that should not contain prompts, graph data or model selections. This affects newly written videos and does not remove metadata from existing files.

Improved audio continuity

Version 3.0 corrects how protected audio overlap and additional history are supplied to continuation scenes.

Extra audio history now occupies only the range immediately before the protected overlap. The same audio positions are no longer passed twice.

For normal video and audio continuity, the recommended starting point remains:

  • Overlap: 22 frames

  • Preserve audio overlap: enabled

  • Extra audio history: 0 frames

A 5-frame overlap is supported as a phase-safe minimum, but it has a higher risk of motion changes, ambience resets or audible scene transitions.

Improved long-audio references

Audio use → Follow scenes · REF2VA automatically advances through a long audio source according to the saved film timeline.

It accounts for:

  • scene duration;

  • protected overlap;

  • regeneration;

  • the saved end of the previous scene;

  • custom Start, End and Extra offset values.

The overlap already contained in the saved latent is not submitted again as a new reference range.

This remains reference guidance for H3. It does not replace the generated soundtrack and cannot guarantee exact reproduction of speech, lyrics, timing or musical rhythm.

Improved seam handling

Repair audio seams can detect and soften isolated electrical clicks near scene joins without shifting samples or changing film duration.

A broad volume drop, restarted ambience or missing generated sound cannot be reconstructed safely during export. LongForge reports these conditions instead of silently rewriting a large part of the soundtrack.

For problematic joins, regenerate the affected continuation with a 22-frame overlap.

Safer projects on Windows

Project handling has been strengthened for Windows Portable installations:

  • per-project operation locking;

  • atomic JSON updates with retry handling;

  • protection against simultaneous rendering and destructive project actions;

  • proper release of Windows file handles after success or failure;

  • safer recovery when a take or project pointer is damaged;

  • continuation after restarting ComfyUI;

  • preservation of older regenerated takes until explicit cleanup.

Clean unused removes takes outside the active continuation only after confirmation. Delete film removes the complete selected project.

Cleaner native workflow

The legacy LongForge LoRA and attention wrapper nodes have been removed from the supplied graph.

Version 3.0 uses:

  • native ComfyUI model, CLIP, VAE and LoRA loaders;

  • Patch Sage Attention KJ;

  • MiniMax H3 Memory Efficient Sage Attention Patch;

  • LongForge nodes only for project, continuity, references, HQ processing and export.

Everything remains visible and manually adjustable. Turbo is optional, LoRAs are optional, and there is no forced sampling preset or 1 MP resolution restriction.

Two generation modes

  • FL2VA — text-to-video, image-to-video and optional first/last-frame guidance.

  • REF2VA — up to 9 images, 3 videos with paired soundtracks, and 3 standalone audio references.

Switch the pipeline inside the same workflow and select the matching MiniMax H3 model.

Resolution, frames, overlap, steps, sampler, scheduler, denoise, CFG, seed and model shifts remain manually adjustable.

Download

Download both archives:

  • ComfyUI-H3-LongForge-NodePack-3.0.zip — required LongForge custom nodes and installation README.

  • H3_LongForge_3.0_Workflow_Guide.zip — unified FL2VA/REF2VA workflow and complete Russian/English user and prompt guide.

Saved LongForge projects and takes from earlier compatible versions can be opened with the v3.0 workflow.

Requirements

Version 3.0 requires:

  • a recent ComfyUI build with native MiniMax H3 support and the V3 node API;

  • a compatible MiniMax H3 FL2VA or REF2VA diffusion model;

  • MiniMax text encoder;

  • MiniMax H3 video VAE;

  • MiniMax H3 audio VAE;

  • ComfyUI-KJNodes for the two optional attention patch nodes;

  • compatible SageAttention when those attention patches are enabled;

  • optional compatible FL2VA/REF2VA LoRAs;

  • an optional compatible 3D latent-upscaler checkpoint for HQ Finish.

MiniMax H3 models:
Comfy-Org / MiniMax-H3

Text encoders:
INT8 ConvRot · NVFP4

Optional latent-upscaler checkpoint:
LBH-123-AI / Minimax H3 Latent Upscaler

Use a compatible minimax_h3_latent_upscaler_3d_conv_v1_*.safetensors checkpoint. The stable LongForge path does not accept the 2D version.

Reference-media loading and latent-upscaler integration are built into LongForge. No separate media-loader or upscaler node pack is required.

Installation on Windows Portable

  1. Close ComfyUI completely.

  2. Back up your existing LongForge folder if needed.

  3. Replace the previous installation with the new folder:

    ComfyUI/custom_nodes/ComfyUI-H3-LongForge-NodePack

  4. Do not merge files from different LongForge versions or keep two installed copies.

  5. From ComfyUI_windows_portable, install the included dependency:

python_embeded\python.exe -m pip install -r ComfyUI\custom_nodes\ComfyUI-H3-LongForge-NodePack\requirements.txt
  1. Install or update ComfyUI-KJNodes if you want to use the supplied attention nodes.

  2. Restart ComfyUI and refresh the browser with Ctrl+F5.

  3. Open H3_LongForge_FL2VA_REF2VA.json.

  4. Select the H3 model, text encoder, video VAE and audio VAE.

  5. Select optional LoRAs before enabling their nodes with Ctrl+B.

Both LoRA loaders, both attention patches and both FL2VA image loaders start bypassed.

Optional HQ Finish installation

To use HQ Finish:

  1. Download a compatible 3D-conv upscaler checkpoint.

  2. Place it in:

    ComfyUI/models/latent_upscale_models/

  3. Restart ComfyUI.

  4. Select the checkpoint in HQ Finish.

  5. Choose UPSCALE or UPSCALE + REFINE.

Keep the same HQ mode, checkpoint, precision and output geometry throughout one film. If you change them, regenerate from Scene 1.

HQ Finish increases generation time, VRAM usage and disk usage because every take keeps the native continuation latent and a separate HQ delivery video.

Generate your film

Write one prompt per scene and use + Scene to add the continuation:

VIDEO: Describe the subject, action, setting and camera.
SOUND: Describe ambience, effects and dialogue.
MUSIC: No music.

These headings are optional. Continuation prompts should describe the next stage of the same action instead of restarting the original scene.

For REF2VA, use native markers such as <Picture 1>, <Video 1> and <Audio 1>, or define aliases such as:

hero=image1
motion=video1
score=audio1
voice=video1_audio

Then use {hero}, {motion}, {score} or {voice} in the scene prompt.

Main actions

  • GENERATE NEXT + ONE SCENE — generate one next scene.

  • GENERATE NEXT + ALL PENDING — generate every remaining scene in order.

  • REGENERATE SELECTED / LAST — rebuild from the selected or last saved scene using its saved seed.

  • VARY SELECTED — regenerate the selected scene with its saved seed plus one.

  • PREVIEW SELECTED — recreate a missing preview without running diffusion.

  • EXPORT FILM — assemble the active saved scenes into one MP4 with generated audio.

When all listed cards are already saved, GENERATE NEXT regenerates the selected scene. Add another card before running if you want to extend the film.

Each successfully generated scene is committed immediately, including during ALL PENDING.

Important HQ behavior

Base H3 generation still runs at the selected native resolution.

After each scene:

  1. the native AV latent is saved for continuation;

  2. a copy of the video latent is upscaled;

  3. optional clean-model refinement is applied;

  4. the HQ scene becomes the preview and export version.

Therefore, a film generated at 736 × 416 with a 2.0× HQ scale is exported at approximately 1472 × 832, while continuation still uses the original native H3 latent.

Optional final post-processing

If you need additional enhancement or frame interpolation after export, you can process the completed film with:

DLSS 5 Visual Enhancer

Download DLSS 5 Visual Enhancer

It is a separate application and is not required by LongForge.

My Telegram channel

AI Bob Public

===========================================================

MiniMax H3 LongForge v2.0 — FL2VA & REF2VA

Support the project

If LongForge helps you, you can support future development through the Tip button on my Civitai profile or an optional donation:

USDT · TRON (TRC20)
THLU3yu8VpoPqP4bMLCp86FRhoCweJ2Qz4

Please send only USDT through the TRON (TRC20) network. Thank you!


One workflow to generate, continue, review and export long videos with audio in ComfyUI.

LongForge transfers saved video and audio latent context between scenes to help preserve motion, appearance, camera movement and ambience. Everything stays in one editable graph with visible loaders and manual controls.

Version 2.0 brings a redesigned scene editor, a built-in reference library, persistent scene previews and automatic slicing of long audio references.

What’s new in v2.0

  • Redesigned Scene Studio: a searchable scene list, clear scene statuses, and an integrated prompt editor and video player.

  • Persistent scene previews: return to earlier fragments after generating additional scenes.

  • Built-in Reference Library: add multiple images, videos and audio files through drag-and-drop or + Add files. Preview, reorder, trim and enable individual sources.

  • Follow scenes audio mode: automatically supply the correct part of a long audio reference to each scene, accounting for overlap, continuation and regeneration.

  • Reference audio switch: disable audio references without muting the generated film.

  • Resolution and duration selectors: convenient choices with Custom controls for manual values.

  • Native model and LoRA loaders: two optional LoRAs, with no forced Turbo mode or fixed step count.

  • Clearer render feedback: generation progress and processing information in Render & Export.

  • Saving and interface fixes: improved project saving and corrected settings shifting after a page reload.

Two generation modes

  • FL2VA — text-to-video, image-to-video and optional first/last-frame guidance.

  • REF2VA — up to 9 images, 3 videos with paired soundtracks, and 3 standalone audio references.

Switch modes inside the same workflow and select the matching H3 model.

Resolution, frames, overlap, steps, sampler, scheduler, CFG and seed remain manually adjustable. There is no 1 MP resolution cap; practical limits depend on your hardware and model.

Long audio references

In Reference Library, add an audio file and choose:

Audio use → Follow scenes · REF2VA

LongForge advances through the source automatically, supplying each scene with its corresponding audio window. There is no need to reconnect files or manually seek between generations.

Start, End and Extra offset · frames control the source range. Follow scenes works throughout the film even when fixed image/video references apply only to the first scene.

This is reference guidance for H3, not a replacement soundtrack. The model generates its own audio and may change the source’s voice, wording, music or rhythm.

Download & installation

Download both archives:

  • ComfyUI-H3-LongForge-NodePack.zip — the custom nodes and installation README.

  • H3_LongForge_FL2VA_REF2VA_PROMPT_GUIDE.zip — the unified workflow and one complete Russian/English user and prompt guide.

  1. Close ComfyUI and extract the node-pack folder into ComfyUI/custom_nodes/. When updating, replace the previous LongForge installation without merging versions. Keep your saved projects in ComfyUI/output/.

  2. Follow the README to install dependencies. The supplied graph uses native MiniMax H3 support, ComfyUI-KJNodes and compatible SageAttention.

  3. Restart ComfyUI, refresh the browser with Ctrl+F5, and open H3_LongForge_FL2VA_REF2VA.json.

  4. Select your diffusion model, text encoder and both VAEs.

  5. Both native LoRA loaders start bypassed. Select a LoRA and press Ctrl+B on its node to enable it when needed. Optional image loaders also start bypassed.

Models: Comfy-Org / MiniMax-H3
Text encoders: INT8 ConvRot · NVFP4
Attention patch nodes: ComfyUI-KJNodes

Reference media loading is included in LongForge; no separate media-loader node pack is required.

Generate your film

Write one prompt per scene, using + Scene to add the next:

VIDEO: Describe the subject, action, setting and camera.
SOUND: Describe ambience, effects and dialogue.
MUSIC: No music.

These headings are optional. In continuation prompts, describe the next stage of the action while preserving the scene’s established appearance and movement.

For REF2VA, use native markers such as <Picture 1>, or define optional names such as hero=image1 and use {hero} in the prompt.

  • GENERATE NEXT + ONE SCENE — generate one next scene per Run.

  • GENERATE NEXT + ALL PENDING — generate all remaining scenes in order.

  • Preview scene — review a selected saved fragment.

  • REGENERATE SELECTED — replace from the selected scene and rebuild its continuation.

  • EXPORT FILM — assemble the active saved scenes into one MP4 with generated audio.

Each successful scene is saved automatically. Reopen your project to continue later.

Optional post-processing

For additional enhancement, upscaling or frame interpolation, you can process the exported film with DLSS 5 Visual Enhancer. It is a separate application and is not required by LongForge.

Download DLSS 5 Visual Enhancer

My Telegram channel

AI Bob Public

===========================================================

MiniMax H3 LongForge v1.0 — FL2VA & REF2VA

Support the project

If LongForge helps you, you can support future development through the Tip button on my Civitai profile or an optional donation:

USDT · TRON (TRC20)

THLU3yu8VpoPqP4bMLCp86FRhoCweJ2Qz4

Please send only USDT through the TRON (TRC20) network. Thank you!

One workflow for generating, continuing and exporting long videos with synchronized audio in ComfyUI.

LongForge carries saved video and audio latent context between scenes to help preserve motion, appearance, camera movement and ambience. Everything stays in one editable graph, with visible model loaders and manual controls.

Two generation modes

  • FL2VA — text-to-video, image-to-video, and optional first/last-frame guidance.

  • REF2VA — image, video and audio references: up to 9 images, 3 videos with paired soundtracks, and 3 standalone audio references.

Switch modes inside the same workflow and select the appropriate H3 model.

Main features

  • One scene editor: add prompts with + Scene.

  • Flexible generation: process one scene or all pending scenes.

  • Automatic continuation: each new scene starts from the saved film end.

  • Saved projects: reopen and continue your film later.

  • Scene regeneration: replace a selected scene and rebuild its continuation.

  • Manual settings: resolution, frames, overlap, steps, sampler, scheduler, CFG and seed.

  • Two optional LoRA loaders: use either, both, or leave them on NONE.

  • Visible attention patches: KJ Sage Attention and MiniMax H3 Memory Efficient Sage Attention.

  • Fragment previews and full-film export in the same graph.

Download & installation

Download both archives:

  • ComfyUI-H3-LongForge-NodePack.zip — required custom nodes and installation README.

  • H3_LongForge_FL2VA_REF2VA_PROMPT_GUIDE.zip — the workflow and one Russian/English prompt guide.

  1. Extract the node-pack folder into ComfyUI/custom_nodes/.

  2. Follow its README to install dependencies. The supplied graph requires native MiniMax H3 support, ComfyUI-KJNodes and compatible SageAttention.

  3. Restart ComfyUI and open H3_LongForge_FL2VA_REF2VA.json.

  4. Select your diffusion model, text encoder, video VAE and audio VAE. Choose optional LoRAs and enable any image loaders you need.

Models DiT: Comfy-Org / MiniMax-H3

Text Encoders: INT8 ConvRot or NVFP4

Other Custom Nodes: ComfyUI-KJNodes — required only for SAGE Attention and MiniMax H3 Memory Efficient Sage Attention.

Generate your film

Write one prompt per scene. A simple structure works well:

VIDEO: Describe the subject, action, setting and camera.
SOUND: Describe ambience, effects and dialogue.
MUSIC: No music.

These headings are optional. The included guide explains both modes, continuation prompts and reference numbering.

In REF2VA, use native markers such as <Picture 1>, or optional names such as {hero}. References can apply to the first scene or every scene.

Choose GENERATE NEXT and press Run to generate. When ready, choose EXPORT FILM and press Run to assemble the saved scenes into one MP4 with audio.

Optional final post-processing

If additional upscaling or frame interpolation is required, process the completed video afterward with:

DLSS 5 Visual Enhancer

It supports RTX Video Super Resolution, DLSS Neural Rendering, and optional frame generation. It is a separate application and is not required for these ComfyUI workflows.

Download DLSS 5 Visual Enhancer

My Telegram channel:

https://t.me/aibobpublic