Sign In

Minimax H3 Dasiwa Xtra - Experimental Optimizations / Functions

Download

1 variant available

Config Other

Dasiwa Xtra 1.5 - Turbo 4-8 Steps.json

668.79 KB

Verified:

Type
ComfyUI Workflows
Stats

159

Reviews
Published

Oct 5, 2026

Base Model

MiniMax H3

Hash
AutoV2
61E8836FFE
default creator card background decoration
Followers - 35

35

Likes - 141

141

Generation, training and LoRA distribution on Civitai are covered by Civitai’s own license agreement with MiniMax. If you download these weights and run them yourself, your use is instead governed by the MiniMax H3 Community License Agreement, whose grant excludes the European Union, the United Kingdom, the Republic of Korea and the United States of America.

MiniMax H3

Dasiwa Xtra 1.5 — MiniMax H3 Workflow

This is an expanded, customized version of the DaSiWa MiniMax H3 workflow, merged with the latest original workflow updates used for this release. It keeps the original MiniMax H3 generation modes and adds workflow controls for attention, acceleration, video continuation, saved-latent continuation, color transfer, and post-processing.

This description covers the updated Dasiwa Xtra 1.5 - Turbo 4-8 Steps.json. The separate VSA experiment is not part of this version and is not required to use it.

What’s new in 1.5

  • Updated DaSiWa base: incorporates the newer original workflow structure and keeps T2VA, I2VA, FL2VA, and REF2VA generation modes.

  • Director-centered prompting: prompt authoring is handled in the MiniMax H3 Director; the old standalone external prompt box is removed. The Director supports ordered image, video, and audio references with per-reference prompts.

  • Unified Settings: attention, acceleration, steps, continuation, color transfer, upscaling, and post-processing controls are grouped and documented in the workflow.

  • Continuation source selector: continue from an uploaded clip or directly from the latest saved MiniMax H3 AV latent.

  • Reduced repeat processing: clip continuation extracts only the final 39-frame tail. Masked AV continuation can reuse its encoded H3 prefix cache when the source file and geometry have not changed.

  • Color Transfer replaces Color Match: use a fixed uploaded image or let clip continuation provide a clean reference frame. Manual image references take priority. Default strength is 0.55.

  • Native / Turbo 4 / Turbo 8 selector: automatically chooses the matching FL2VA or REF2VA Turbo LoRA, denoising step count, and video/audio shifts. Native keeps manual steps and shifts. No PDD or VSA branch is included in this version.

  • Corrected sigma scheduling: generation schedulers and guiders now receive the same SigmaShift-patched model, so the noise schedule follows the selected preset instead of the pre-shift model.

  • Independent acceleration switches: Spectrum and Sparse Attention follow their Settings switches in Native, Turbo 4, and Turbo 8. Sparse supports Native, SageAttention, or Kitchen with dense fallback; legacy Sol-Attn remains excluded.

  • General sampling recommendation: sa_solver + simple is the recommended starting combination for Native, Turbo 4, and Turbo 8. Benchmark results are scene-specific, not a guarantee of identical quality across modes.

  • Repaired latent-upscale branch: corrected the MMH3 branch-control targets, added temporal-parameter control, restored normal/refined latent selection, and routed the upscaler and its scheduler through the matching shifted model. This optional branch remains disabled by default; execution testing is pending after installing its node pack.

  • Expanded finishing tools: optional LBH latent upscale and four-step refinement, “Upscale Last Job,” tiled H3 AV upscale, image upscaling, RIFE interpolation, watermarking, and fast H3 preview.

Generation and prompting

The workflow supports the MiniMax H3 Director modes T2VA, I2VA, FL2VA, and REF2VA. Use the Director’s mode-specific prompt format and connect its guide output to Settings. The built-in examples and Director prompting guide show the expected prompt structure.

Director-provided width, height, frame rate, and duration feed the workflow automatically. The workflow includes aspect-ratio and resolution controls for choosing an output size while keeping dimensions aligned with MiniMax H3 requirements.

Attention and acceleration

The Attention selector chooses one backend:

  • Value: 0 · Backend: Native · Notes: Best compatibility; useful as a baseline.

  • Value: 1 · Backend: SageAttention · Notes: Requires a compatible SageAttention installation.

  • Value: 2 · Backend: Kitchen · Notes: Recommended starting point for speed, quality, and compatibility.

  • Value: 3 · Backend: Sol-Attn · Notes: Optional experimental backend; requires ComfyUI-SolAttn_triton.

Sparse Attention is a separate toggle. It supports Native, SageAttention, and Kitchen, and is not applied when Sol-Attn is selected. Dense attention remains available as a fallback.

The Acceleration selector provides:

  • Value: 0 · Mode: Native · Steps: Manual Settings value · REF2VA video / audio shift: Manual Settings values · FL2VA video / audio shift: Manual Settings values

  • Value: 1 · Mode: Turbo 4 · Steps: 4 · REF2VA video / audio shift: 12 / 3 · FL2VA video / audio shift: 6 / 3

  • Value: 2 · Mode: Turbo 8 · Steps: 8 · REF2VA video / audio shift: 6 / 3 · FL2VA video / audio shift: 12 / 3

Turbo selects one matching LoRA at strength 1.0. Other enabled LoRAs in the main loader still apply, but do not enable an additional Turbo LoRA there: the selector already supplies it. The current REF2VA 4/8-step and FL2VA 4-step presets use the installed LoRA files listed below. The selected FL2VA 8-step file is the older 544p-targeted variant, not the newer 768p variant; choose the Director resolution accordingly.

Both the scheduler and the guider use the same model after the selected video/audio shifts are applied. The optional MMH3 upscale branch also uses the corresponding shifted model for refinement and sigma scheduling.

Start with sa_solver + simple for Native, Turbo 4, or Turbo 8. This is the workflow's general recommendation; the documented artifact-free SA Solver result was measured with Turbo 4, and equivalent results have not been confirmed in every mode.

In our scene, res_multistep + simple gave excellent Native 20-step results, but showed visual artifacts with Turbo 4 even when Spectrum and Sparse were OFF and Attention was Native. Switching to sa_solver + simple produced a Turbo 4 result with no reported visual artifacts. This points to the sampling configuration rather than proving that the Turbo LoRA alone causes artifacts.

Combining Sparse Attention and Spectrum

Both switches now work independently of the Acceleration selector. You can test Spectrum alone, Sparse alone, or both together with Native, Turbo 4, or Turbo 8. For Sparse, use Attention 0, 1, or 2; selecting legacy Sol-Attn (3) disables the sparse route. Spectrum is no longer automatically bypassed in Turbo.

For a controlled comparison:

  1. Fix the seed, prompt, references, resolution, duration, sampler, and scheduler.

  2. Start with Spectrum and Sparse OFF.

  3. Enable Spectrum alone, then Sparse alone, then both.

  4. Inspect quality, anatomy, motion, camera/framing development, and prompt adherence—not only render time.

BlockSparseAttention: log messages indicate whether the sparse path runs or falls back to dense attention. Dense attention at the start of sampling or in protected transformer blocks is expected; our preset retains dense blocks 0–2 and starts sparse attention at 20% of the sampling window.

Sparse + Spectrum + Kitchen was promising in our Native 20-step scene. Turbo combinations remain optional experiments: disable either switch if ghosting, reduced detail, unstable motion, or other artifacts appear. A successful render does not by itself prove that every optimization executed. Sparse + Spectrum with sa_solver + simple in Turbo still needs its own benchmark.

Local benchmark highlights

  • Configuration: Native 20, RES Multistep / Simple, no acceleration patches · Time: 993 s · Reported result: Excellent quality and motion

  • Configuration: Native 20, RES Multistep / Simple, Sparse + Spectrum + Kitchen · Time: 351.81 s · Reported result: Quite good quality and motion

  • Configuration: Turbo 4, SA Solver / Simple, Native attention, Spectrum and Sparse OFF · Time: 162.34 s · Reported result: No visual artifacts reported

These are individual local runs at Native ShortEdge 768px, not a universal performance ranking. Seed/cache equality was not independently verified across all runs. The updated Spectrum/Sparse controls and repaired MMH3 upscale wiring still need execution testing in this exact workflow file. Full observations are recorded in Dasiwa Xtra 1.5 Benchmark.md.

Continue from clip or last saved latent

Set Continue Source to choose the source:

  • 0 · Clip: choose a video in the optional Load continuation clip node. The workflow extracts only the final 39-frame audio/video tail. Native AddGuide uses a selectable context length; Masked AV is intended for precise audio/video boundary continuation. The two clip continuation modes are routed so they are not applied together.

  • 1 · Last Saved Latent: directly resumes the latest H3 audio/video latent checkpoint created by the workflow. Run a generation once to create a checkpoint first. This path automatically uses Masked AV and avoids decoding the previous output and encoding it through the VAEs again. Keep the same output resolution and compatible H3 model and VAEs.

For clip-based Masked AV continuation, the encoded tail can be cached and reused when the source file and geometry remain unchanged. AddGuide context values of 1, 6, 12, 22, and 39 frames range from minimal context to stronger continuity; 22 is a good starting point.

When not continuing a clip, leave the small internal fallback video connected to Settings. The external continuation video is optional and can remain bypassed.

Color Transfer

Color Transfer applies one uniform Lab color transform before upscale. It can use a clean frame from the continuation clip or an uploaded Color Transfer Reference image. An uploaded manual image takes priority over the clip, which is useful for keeping a fixed look across multiple continuations. When neither source is available, transfer is skipped.

The image loader starts bypassed so the workflow remains usable without a color reference. Un-bypass it and upload an image when you want to use a manual reference, then enable Color Transfer in Settings. The default Color Transfer Strength is 0.55; lower values produce a subtler transfer. Reuse the same original reference across successive clips to reduce color drift.

Upscaling and post-processing

  • LBH Current + 4-Step Refine: optionally generates at base resolution, upscales the latent, then performs four high-resolution refinement steps. For example, 16 total steps become 12 base-resolution steps plus the final four refinement steps. It is compute- and VRAM-intensive; start around 1.25× or 1.50×.

  • Upscale Last Job: loads the latest saved latent checkpoint for the LBH upscale/refinement path. It requires a previous saved checkpoint and is bypassed while Continue Source is set to Last Saved Latent.

  • Latent Upscaler 2x / MMH3 Ultimate Upscale: optional model-based tiled H3 audio/video latent upscaling. The Settings switch controls the parameter, conditioning, scheduling, and upscale nodes; normal/refined output selectors route the correct latent. Install Comfyui-MMH3-UltimateUpscale and its latent-upscale model, restart ComfyUI, and leave this setting OFF unless you intend to run the upscale. Expect substantially longer processing; the repaired branch is not yet execution-validated in this release.

  • Image finishing: optional RIFE 2× frame interpolation, simple or model-based image upscaling, NVIDIA RTX upscale/refine, and watermark overlay.

  • Preview: KJNodes fast preview can use the MiniMax H3 TAE model.

Optional finishing paths increase processing time and may need substantial VRAM or system memory. Keep them disabled for baseline tests, then add one at a time.

Requirements

Use ComfyUI v0.37.0 or newer for the native MiniMax H3, Math Expression, Native AddGuide, and Block Sparse Attention nodes. FFmpeg is required for video/audio muxing and output processing.

Required custom node packs:

Optional attention backends: SageAttention and ComfyUI-SolAttn_triton, when selecting their respective attention options.

Required MiniMax H3 models

Install the files in the indicated ComfyUI model folders:

Optional acceleration and finishing models

Use the exact ComfyUI-format filenames selected by the workflow. The official LightX2V Turbo model repository also provides these Turbo variants; this workflow does not automatically switch to newer releases.

Results and performance depend on GPU/VRAM, ComfyUI and PyTorch versions, model precision, resolution, frame count, and enabled acceleration or post-processing options. Start with Native acceleration and conservative finishing settings, then benchmark changes individually.