Sign In

MacMax Turbo V1.2 (4-Step Turbo LoRA)

Updated: Sep 23, 2026

toolapple siliconminimax h3

Download

1 variant available

Archive Other

macmax-minimax-h3-mac.zip

50.03 KB

Verified:

Type
ComfyUI Workflows
Stats

336

Reviews
Published

Aug 7, 2026

Base Model

MiniMax H3

Hash
AutoV2
F33C83DE0B
default creator card background decoration
Followers - 25

25

Likes - 77

77

Generation, training and LoRA distribution on Civitai are covered by Civitai’s own license agreement with MiniMax. If you download these weights and run them yourself, your use is instead governed by the MiniMax H3 Community License Agreement, whose grant excludes the European Union, the United Kingdom, the Republic of Korea and the United States of America.

MiniMax H3

v1.2 adds an optional turbo-LoRA fast path: 4 steps for silent b-roll, 6 steps for anything with audio (dialogue, foley, ambience). Needs the pruned GGUF DiT. Base 20-25 steps is the quality lane. Full notes in the README.

MacMax runs MiniMax H3 locally on a Mac. Video and stereo audio come out of the same pass, so foley and dialogue are generated with the picture, not layered on after.

MacMax runs text to video, image to video and first/last frame in one graph, each with native audio.

Measured on a 48 GB M5 Pro, ComfyUI 0.30.0. Full docs, every measurement: https://github.com/Bambushu/minimax-h3-mac

Render times

0.5 MP vertical, Base int8 path; the turbo GGUF path drops to 4 steps silent / 6 with audio.

base, 20 steps:
3s  image to video                            ~14 min
5s  text to video                             ~24 min
5s  image to video, chained link              ~39 min

turbo, 4 silent / 6 with audio:
3s  silent, 4 steps                            ~7 min
4s  spoken, 6 steps                           ~11 min

Cost tracks megapixels x seconds. A first-frame image adds little; duration and resolution are the levers.

Sizing: about 22k tokens is comfortable on 48 GB. That is 0.6 MP at 5s, or 1.03 MP at 3s.

Chaining

Clips continue each other: motion carries across the cut, the scene holds, and so does the audio bed. Un-bypass Save Latent on a clip you may want to continue and it writes a small latent; to continue it, un-bypass three more nodes, set two clip indexes, queue.

The continuation comes back exactly 22 frames shorter, because those frames are the pinned context and they get trimmed so the files concatenate cleanly. If yours is not 22 frames shorter, chaining did not engage. That is the check worth doing.

Nine links ran as one sequence, 39s of continuous scene. The 25s video in the gallery is the front of it, hard concatenated, no crossfades and no level matching, so every join is visible as rendered.

Two rules that cost me renders:

  • Write each beat to fill the whole clip. If the action finishes early, the model can fill the rest by cutting to an animated version of one of your reference images. Reseeding does not fix it.

  • Clip 1's framing is inherited by everything after it. Lock it before you start.

Setup

Four model files, about 41 GB, all linked in the README. Use the GGUF text encoder, the stock one is CUDA only.

ComfyUI 0.30.0 in its own checkout, launched with:

ASFP8_INT8_EXT=1 python main.py --port 8288 --reserve-vram 10 --cache-none --disable-smart-memory

Three node packs are required and two more are optional, for chaining and the smaller encoder. ./install_node_packs.sh all clones the lot. Everything the optional packs add ships bypassed, so neither is needed to render.

48 GB is what this was measured on. 32 GB works too, reported by users rather than tested here; expect to stay at the shorter durations.