Updated: Aug 10, 2026
base modelDownload
6 variants available
bf16 SafeTensor
minimax_h3_ref2va_bf16.safetensors
BF16, good balance • 61.73 GB
Unverified: Scan requested
SafeTensor
int8
minimax_h3_ref2va_pruned_int8_convrot.safetensors
8-bit integer, smaller file
Verified: 6 days ago
You need these files to run this model. We'll show the best match for your preferences.
VAE
minimax_h3_video_vae_fp16.safetensors selected • 4.85 GB
Downloads your preferred variants
minimax_h3_fl2v_lightx2v_turbo_4step_v0.1_comfy.safetensors
.safetensors • 1.82 GB
Verified: 3 days ago
minimax_h3_fl2v_lightx2v_turbo_4step_v0.1_comfy_resized_avg_rank_21_bf16.safetensors
.safetensors • 300.29 MB
Verified: 3 days ago
minimax_h3_fl2v_lightx2v_turbo_4step_v0.1_comfy_resized_avg_rank_21_bf16.safetensors
.safetensors • 300.29 MB
Verified: 3 days ago
1,0120 1 2 3 4 5 6 7 8 9,0 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 9
(96)
Aug 10, 2026
MiniMax H3

17.3K0 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 9.0 1 2 3 4 5 6 7 8 9K
1.4M0 1 2 3 4 5 6 7 8 9.0 1 2 3 4 5 6 7 8 9M
9.8M0 1 2 3 4 5 6 7 8 9.0 1 2 3 4 5 6 7 8 9M
License:
Hailuo H3 is the newest video generation model from MiniMax, the Shanghai AI lab behind the Hailuo video app and the MiniMax text and audio model families. H3 generates native 2K video with synchronized audio in a single pass, at clip lengths up to 15 seconds.
Developed by MiniMax. All credit for the model goes to the MiniMax team. This is a hosted integration, not a weight mirror - H3 is a closed model served through the MiniMax API, and Civitai runs your generations against it so you can work on-site alongside the rest of your library. There are no weights to download here or anywhere else.
Background
MiniMax has shipped video models under the Hailuo name since 2024, moving from the I2V-01 series through Hailuo 02 and Hailuo 2.3. Those earlier models topped out around 1080p with fixed 6 or 10 second clips and no native sound. H3 is the step past that: it raises the ceiling to 2K, opens the clip length to any integer from 4 to 15 seconds, and folds audio generation into the same pass as the picture rather than bolting it on afterward.
Capabilities
• Native 2K output (2560x1440 at 16:9)
• Clips from 4 to 15 seconds, in one second increments
• Synchronized audio generated with the video - dialogue, effects, and ambience timed to what is on screen
• Text-to-video, image-to-video, and first frame plus last frame control
• Omni-reference conditioning: up to 9 reference images, 3 reference video clips, and 3 reference audio clips
• Multi-shot sequences and instruction-based editing
On Civitai
H3 is wired into the video generator across four workflows: Create Video (text only), Image to Video (single first frame), First/Last Frame (both endpoints), and Reference to Video (up to 9 reference images). Resolution is fixed at 2K because that is the only output H3 produces. Aspect ratio is selectable at 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16 for text prompts; when you supply a frame image, H3 adapts the framing to that image instead.
Reference video and reference audio inputs are supported by the model but are not exposed in the Civitai form yet.
Notes
H3 is a new model and MiniMax has not published a full spec sheet or a stable pricing page for it. Behavior and cost may shift as they finish the rollout. Prompt handling follows the usual MiniMax posture - this is a commercially hosted service and content is subject to the provider's policies, so expect refusals on material their API declines.
Links
• MiniMax video generation docs
• Hailuo AI
• MiniMax
• MiniMax Hailuo AI Terms of Service