Sign In

MiniMax-H3 Ref2VA (4-Bit INT4 Safetensors)

Download

1 variant available

int4 SafeTensor

MiniMax-H3-Ref2VA-INT4.safetensors

4-bit integer, smallest β€’ 9.8 GB

Verified:

Type
Checkpoint Trained
Stats

21

Reviews
Published

Aug 22, 2026

Base Model

Other

Hash
AutoV2
F302A1612F
default creator card background decoration
Followers - 423

423

Likes - 1296

1.3K

Text-Tacular Contest Winner

License:

github-social-1280x640.png

# ⚑ MiniMax-H3 Ref2VA (Safetensors Quantized Checkpoints)

This repository provides optimized Safetensors quantizations of MiniMax-H3-Ref2VA (Reference-to-Video & Audio), engineered specifically for consumer GPUs (16 GB VRAM, RTX 4080 / RTX 4090 / RTX 4090 Mobile / RTX 3090).

Choose between two purpose-built variants depending on your workflow and VRAM budget:

---

## 🌟 Variant 1: Digital Forge Turbo W4A8 (All-In-One 4-Step) β€” Recommended

The flagship mixed-precision release designed by Digital Forge (DF). It eliminates common INT4 artifacts while mathematically fusing the official LightX2V 4-Step Turbo adapter directly into the model weights.

### πŸ›‘οΈ Why Choose DF-Turbo-W4A8?

πŸš€ *Pre-Baked 4-Step Turbo**: The 4-step Turbo LoRA is fused before quantization. No external LoRA loading requiredβ€”saves 1.28 GB runtime VRAM and eliminates 624 GEMMs per step!

πŸ›‘οΈ *Zero "Zombie Sway" (Pristine BF16 AdaLN)**: All 335 time-modulation layers, LayerNorms, and 1,025 calibration grid points remain in unquantized BF16, ensuring continuous temporal velocity ($\vec{v}$) and smooth, natural character motion.

🎯 *Zero Reference Feature Drift (INT8 Cross-Attention)**: Uses INT8 per-channel quantization across all cross-attention to_k and to_v layers to preserve fine facial features, eye contact, and clothing textures without outlier clipping.

πŸ’Ύ *16 GB VRAM Resident (12.62 GiB)**: Fits 100% resident inside 16 GB GPUs with zero PCIe swapping.

### βš™οΈ Recommended Settings (DF-Turbo-W4A8)

* Sampling Steps: 4

* CFG Guidance: 1.0 (Distilled trajectory)

* Video Shift: 12.0

* Audio Shift: 3.0 (32 kHz synchronized audio)

* Frame Count: 25 to 124 frames

---

## πŸ§ͺ Variant 2: Pure INT4 (Ultra-Low VRAM Experimental)

A pure 4-bit uniform quantization designed for minimal memory consumption.

### πŸ”¬ Model Details & Caveats

* VRAM Footprint: *9.80 GiB** (Ultra-compact, maximum memory headroom).

* Quantization Scheme: Symmetric 4-Bit Linear FastInt4Linear) with Group-128 scaling.

* Adapter Support: Requires loading the external LightX2V 4-step Turbo LoRA minimax_h3_ref2v_turbo_4step_v0.1_bf16.safetensors) if 4-step generation is desired.

⚠️ *Experimental Note**: In extended runs (e.g. 10-second / 124+ frame sequences), pure INT4 quantization on time-modulation layers may exhibit slight identity drift or rhythmic "zombie sway" artifacts. Use DF-Turbo-W4A8 if artifact-free motion is required.

---

## ⚑ Performance & Engine Compatibility (Both Models)

* Zero-Copy Loading: Loads via OS virtual memory mapping mmap) in ~0.03 to 0.05 seconds.

* TeaCache Compatible: Full native support for Block-Level TeaCache (~50% compute step bypass).

* Attention Acceleration: Supports SageAttention 2.0 Patched (Fast FP8 with FP32 accumulator), FlashAttention-2, and native PyTorch SDPA.

* VAE Tiling: Recommended VAE tile size of 512 for efficient 3D temporal decoding.

---

## πŸ“œ License & Attribution

Base model licensed under the *MiniMax H3 Community License Agreement**, Copyright Β© 2026 MiniMax.

Powered by *MiniMax H3** & Digital Forge.