Updated: Sep 3, 2026
base modelDownload
3 variants available
fp16 SafeTensor
minimax_music3_dit_fp16.safetensors
Half precision, best balance • 4.58 GB
Verified: 21 days ago
SafeTensor
You need these files to run this model. We'll show the best match for your preferences.
Text Encoder
minimax_music3_text_encoder_pruned_bf16.safetensors selected • 15.56 GB
bf16
minimax_music3_text_encoder_pruned_bf16.safetensors
BF16, good balance
Verified: 21 days ago
int8
minimax_music3_text_encoder_pruned_int8_convrot.safetensors
8-bit integer, smaller file
Verified: 21 days ago
Downloads your preferred variants
260 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 9
00 1 2 3 4 5 6 7 8 9
(5)
Sep 3, 2026
MiniMax Music 3

17.5K0 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 9.0 1 2 3 4 5 6 7 8 9K
1.4M0 1 2 3 4 5 6 7 8 9.0 1 2 3 4 5 6 7 8 9M
10.8M0 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 9.0 1 2 3 4 5 6 7 8 9M
License:
MiniMax-Music3 Community LicenseMiniMax Music 3 is a high-performance music generation model for creating complete songs up to five minutes long. Conditioned on lyrics and a detailed music description, it generates structurally coherent songs with expressive vocals, evolving arrangements, and stable long-form audio quality.
MiniMax Music 3 combines an 8B Global LLM for long-range musical structure, a 0.6B Local LLM for frame-level acoustic detail, and a continuous hidden-state synthesis system based on Flow Matching and Flow-VAE. The model produces 32 kHz, 16-bit stereo WAV audio.
Complete Songs with Long-Range Coherence
MiniMax Music 3 natively supports full-song generation up to five minutes. The model maintains musical themes, rhythm, vocal identity, and arrangement progression across long sequences, enabling complete structures such as intro, verse, pre-chorus, chorus, bridge, instrumental break, and outro.
Fine-Grained Music Control
The model accepts two complementary inputs:
Lyrics define the words to be sung and may include explicit section tags such as
[Intro],[Verse],[Pre-Chorus],[Chorus],[Post-Chorus],[Bridge],[Instrumental],[Solo], and[Outro].Music description defines the musical style, emotional progression, vocal performance, instrumentation, arrangement, and production profile.
For precise control, we recommend using a Structured Caption with three sections:
Global Metadata: genre, subgenre, BPM, key, scale, emotional progression, listening scenario, and production profile.
Vocal Details: vocal gender, timbre, performance style, harmony, backing vocals, and vocal effects.
Arrangement: primary and secondary instruments, section-level instrument evolution, groove, bass, percussion, textures, and spatial effects.
This representation allows the model to follow not only a global style, but also the musical development of the song over time.
