Download
1 variant available
bf16 SafeTensor
minimax_h3_hybrid_pruned_BF16.safetensors
BF16, good balance • 37.54 GB
Verified: 13 days ago
880 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 9
(9)
Sep 6, 2026
MiniMax H3
This is for experimental purposes. I cannot accept responsibility for any consequences arising from its use.

2740 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 9
2700 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 9
MiniMax H3 is licensed by MiniMax under the MiniMax H3 Community License Agreement. That agreement’s Applicable Territory excludes the European Union, the United Kingdom, the Republic of Korea and the United States of America. Your use of H3 and of any H3 derivative is subject to that agreement and its Acceptable Use Policy.
MiniMax H3
【MiniMax H3_SparseRef15_Hybrid】
MiniMax H3_SparseRef15_Hybrid is an experimental of Fl2VA-based hybrid models family that incorporates selected strengths of Ref2VA without any LoRA merging. It is designed to improve reference consistency while preserving natural motion, scene flexibility, and long-form stability.
Note: This is not a native Ref2V model. It is based on Fl2VA with an added "Ref" effect, so please use LoRAs designed for Fl2VA rather than Ref2V LoRAs.
The core concept is a sparse Ref2VA influence applied through AdaLN across 15 main transformer blocks, rather than using Ref2VA uniformly throughout the model.
This sparse structure is intended to retain more of FL2VA's motion freedom and scene behavior while strengthening character identity and reference consistency.
In testing, the model has shown strong resistance to character drift across long multi-clip generations, including cases where the character changes direction, temporarily leaves a clear frontal view, or continues through many consecutive clips.
Different versions may vary in model size, precision, quantization, speed, memory usage, visual quality, and reference strength.
For reference, Pruned_INT8 is a standard, lightweight pruned model in which the full model's "time-conditioning" has been compressed to 8 dimensions and the large-scale Attention/MLP matrices have been converted to the INT8 ConvRot format.
In contrast, the Pruned_Partial-INT8 series also converts large-scale Attention/MLP matrices to the INT8 ConvRot format but reconstructs and retains critical components—such as AdaLN and time-conditioning—at FP32 precision.
<Hybrid_Pruned_INT8>
A lightweight Pruned Partial-INT8 version of SparseRef15 Hybrid.
Technical characteristics:
- 15 sparsely distributed Ref2VA AdaLN blocks within the 50-block main transformer
- Pruned H3 architecture
- Partial INT8 quantization
- Designed for lower memory usage and faster inference
- Strong emphasis on character identity consistency and long-form stability
In extended multi-clip testing, this version maintained character identity with very little visible drift, even over long sequences.
On suitable hardware, it is considerably faster and lighter than the heavier non-INT8 variant while retaining good motion quality and overall visual coherence.
<Hybrid_Pruned_Partial-INT8 Ver.1.0>
A newer Pruned Partial-INT8 Hybrid variant focused on stronger reference consistency and long-form character stability.
Technical characteristics:
- Pruned H3 architecture
- Partial INT8 quantization for reduced model footprint and faster inference
- Hybrid FL2VA / Ref2VA structure
- Reference behavior tuned for stronger character identity persistence
- Designed to remain stable across long multi-clip generations
Compared with "Hybrid_pruned_int8", this version shows noticeably stronger reference persistence.
In long-form testing, character identity remained highly stable across a 21-clip sequence with very little visible drift, even through repeated changes in pose, direction, and clip transitions.
The stronger reference behavior can sometimes reduce scene freedom or resist situations where the character is expected to remain fully hidden for an extended period. In return, this version is particularly well suited to long-form generations where character consistency is the highest priority.
The Partial-INT8 structure also gives this version a much smaller model footprint and significantly faster inference than the heavier non-INT8 Pruned variant on supported hardware.
<Hybrid_Pruned_BF16>
This is a hybrid derived from BF16 that has undergone only minimal pruning.
The hybrid method is the same as Pruned_INT8.