Download
1 variant available
int4 SafeTensor
minimax_h3_ref2va_8steps_quantfunc_int4_r128.safetensors
4-bit integer, smallest • 11.52 GB
Verified: 4 days ago

7
23
Generation, training and LoRA distribution on Civitai are covered by Civitai’s own license agreement with MiniMax. If you download these weights and run them yourself, your use is instead governed by the MiniMax H3 Community License Agreement, whose grant excludes the European Union, the United Kingdom, the Republic of Korea and the United States of America.
MiniMax H3
MiniMax H3 — QuantFunc A4W4 INT4
True 4-bit inference — A4W4 (4-bit activations × 4-bit weights).
QuantFunc's core quantized matrix multiplications use INT4 activations and INT4 weights on its INT4 inference backend. A4W4 describes the compute precision used during inference, alongside the reduced weight size and memory bandwidth demand.
QuantFunc A4W4 uses 4-bit weights and 4-bit activations for MiniMax H3's core quantized matrix computations. Character detail, style fidelity and fast motion stay clear and coherent, while H3's native video+audio generation is fully preserved.
In our internal FL2VA evaluation, QuantFunc INT4 vs the BF16 baseline (same prompt, same seed) measures ~23.7 dB PSNR.
Showcase
All clips below were generated by MiniMax-H3-QuantFunc-4bit — click a player to watch.
Live-action performance Animated chase
Watch generated video with audio

Watch generated video with audio High-speed car chase Stylized character

Watch generated video with audio

Watch generated video with audio
Poster frames are the first frame of each clip. Showcase spec: 896 × 1184, 124 frames, 24 FPS, ~5s.
Same 124 frames, up to 3.19x FP8's per-step speed
On an RTX 4090, 768 × 768, 5s, 124 frames:
Backend Per-step time Relative speed QuantFunc INT4 3.2s — INT8 ConvRot 8.5s 2.66x faster FP8 10.2s 3.19x fasterPer-step core-model inference time only — excludes text encoder, VAE, audio processing and video saving. Actual speed varies with resolution, frame count, driver and software version.
Swap one loader, keep the rest of your workflow
- Install or update ComfyUI-QuantFunc.
- Download the FL2VA 4-step or Ref2VA 8-step weights.
- Swap your model loader for the QuantFunc loader and pick the matching weight file.
Prompts, reference images/videos, audio assets and the rest of your workflow nodes stay unchanged.
Both weight sets already have the acceleration LoRA, Token Refiner, INT4 Refiner and INT8 Conv Sidecar fused in — no extra components to attach.
RTX 20-series through GB300, one build covers it all
Runs on every NVIDIA SM75+ GPU: RTX 20/30/40/50-series, A100, H100, H200, B100, B200, GB300.
4-bit weights significantly cut the weight-bandwidth cost of loading and inference, making MiniMax H3 much easier to run on consumer GPUs. Actual VRAM needs depend on resolution, frame count, reference-asset count and the rest of your workflow.
Choose a model
Variant File Size Recommended steps Use case FL2VAminimax_h3_fl2va_4step_quantfunc_int4_r128.safetensors
12.37 GB (11.52 GiB)
4 steps
Text-to-audio/video, first-frame, last-frame, first+last-frame control
Ref2VA
minimax_h3_ref2va_8steps_quantfunc_int4_r128.safetensors
12.37 GB (11.52 GiB)
8 steps
Multimodal reference generation from images, video and audio
FL2VA supports zero, one or two input images; Ref2VA targets more complex multimodal reference scenarios. See the official MiniMax H3 repo for details on both modes.
Use the matching scheduler for the model. Both weight sets already carry the acceleration LoRA fused in — no need to load it separately.
Loading
Weights use QuantFunc's own sealed safetensors format (quantization parameters and metadata are sealed).
Load with ComfyUI-QuantFunc or the QuantFunc inference engine — this is not a drop-in Diffusers checkpoint.
If the model isn't recognized or fails to load, update ComfyUI-QuantFunc to the latest version and restart ComfyUI.
Technical details
- INT4 weights × INT4 activations (W4A4)
- SVDQuant rank 128, group size 64
- INT4 Token Refiner + INT4 Refiner
- INT8 Conv Sidecar
- FL2VA measured max relative output error from fp16 adaLN folding: ~7.58e-4
- Minimum GPU architecture: NVIDIA SM75
Source & license
This repository distributes derived quantized weights produced from MiniMax H3.
MiniMax H3 and its derivative weights follow the MiniMax H3 Community License Agreement. QuantFunc's quantization tooling and implementation follow their own respective licenses. Please read and comply with the original model's license terms before use.
Official QuantFunc links
- Hugging Face — QuantFunc
- ModelScope — QuantFunc
- Official Discord
- QuantFunc website
- ComfyUI-QuantFunc plugin and workflows
License terms
The MiniMax H3 Community License covers these derived weights. Its standard grant excludes the European Union, United Kingdom, Republic of Korea and United States. Separate authorization is required where the community grant does not apply. Commercial products/services above the license's revenue threshold also require prior written authorization. Redistribution must include the license agreement and its required NOTICE. Full license.
Source and weight integrity
Original QuantFunc release: QuantFunc/Minimax-H3-Quantfunc-4bit.
minimax_h3_fl2va_4step_quantfunc_int4_r128.safetensors— SHA256:fa9526b0891b455e63e293efc52331a82bc69cd3a3787625b49ccbd8145c5544minimax_h3_ref2va_8steps_quantfunc_int4_r128.safetensors— SHA256:053d6d4e7d85ee18d0f0784eb1e28774ed0fef8a1f3175a8dd451eff2254b258
Showcase media is reproduced from the original QuantFunc model card. Exact seeds, prompts and rank variants are not supplied for every example. Performance figures are QuantFunc's reported measurements under the stated conditions; they are not guarantees for other workflows.
