Download
1 variant available
fp32 SafeTensor
manga_tone_v2.safetensors
Full precision, largest file • 591.55 MB
Verified: 8 hours ago
950 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 9
00 1 2 3 4 5 6 7 8 9
(12)
Sep 11, 2026
MiniMax H3

10 1 2 3 4 5 6 7 8 9
120 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 9
MiniMax H3 is licensed by MiniMax under the MiniMax H3 Community License Agreement. That agreement’s Applicable Territory excludes the European Union, the United Kingdom, the Republic of Korea and the United States of America. Your use of H3 and of any H3 derivative is subject to that agreement and its Acceptable Use Policy.
MiniMax H3
A LoRA that teaches MiniMax H3 (video + audio joint DiT) the manga
screentone rendering style with motion: monochrome ink rendering, proper
tone shading, and tone that holds up under dynamic camera movement and action.
Usage
- Trigger phrase: manga-tone rendering
- Recommended strength: 1.0 (usable 0.8 - 1.2)
- Standalone: res_multistep / simple, 14 steps
- With turbo LoRA: euler / beta, 8 steps (ref2v turbo 8-step) or 4 steps
(fl2v turbo 4-step), audio sigma shift 5
- Checkpoints: int8 and mixed int4/int8 pruned convrot conversions
- Load through any H3 LoRA loader (H3LoraStack / LoraLoaderModelOnly)
Example prompt:
manga-tone rendering, a samurai draws his sword in a bamboo grove,
dramatic low angle, speed lines
What it does
Without the LoRA, H3 renders "manga" prompts as grayscale anime with flat
shading and little motion. With v2, outputs are monochrome with ink/midtone
statistics inside the trained manga distribution, and motion is preserved
under action prompts (measured +35% inter-frame motion vs v1 on a
motion-heavy prompt, tone metrics unchanged).
Training
All training data is synthetic; no published or third-party artwork was used:
- 714 static 5-frame crop clips from 58 AI-generated fictional manga pages,
two crop scales
- 80 animated clips: still crops animated by H3 itself (v1 @0.7 + turbo),
then cleaned frame-by-frame with a deterministic monochrome render pass
- 6 windowed clips from two action reference videos produced by feeding the
same fictional manga pages to Google Omni (x4 repeat)
- LoRA rank 32 / alpha 32 on 208 projections, flow matching on the video
stream, lr 5e-5 constant with warmup, 2000 steps, bf16, gradient
checkpointing; v2 initialized from v1
- Audio stream masked in the loss; audio generation unchanged to first
approximation but not guaranteed
Notes
- Tone dot frequency is learned relative to page scale (renders ~9 px period
at 960 px canvas width).
- Abstract motion-graphics tropes (layer decomposition, graphic animation)
remain out of distribution for the base model; action/cinematic prompts
play to this LoRA's strengths.
