Sign In

Training H3 Loras on Small Datasets: What I've Learned So Far

10

Training H3 Loras on Small Datasets: What I've Learned So Far

Minimax H3 is the hot new base model on CivitAI. There are new models all the time but this one interests me because it's the first video model I can train loras for on my own hardware. It reminds me of the early SD1.5 days: its janky, but exciting to experiment with. Hearing a character's voice to go along with a video is wild. It feels like a whole new world has opened up.

I don't have a perfect guide, but I do have a handful of example loras and the datasets I used. I'm sharing them in the hope that you'll share what you've learned too. Please comment your own tips, settings, and/or failures in the comments!

TLDR

- Small datasets work. You don't need a huge, perfect dataset to get started (mine range from 1 to 100 files).

- You can train with images only, images plus audio, video with audio, or a mix of images and videos.

- If you train on CivitAI, use the defaults and never go below 3000 steps.

- Locally, I use Fizgig with its recommended defaults, I only change Dim/Alpha 32, Resolution 1024, and number of epochs.

What you can train

1. Images only

Good for getting a character or style into H3 with an easy dataset.

- https://civitai.red/models/2980244/lavender-goth-citron-oc-h3

- https://civitai.red/models/2943813/minimalistic-style-minimax-h3

- https://civitai.red/models/2943717/botw-zelda-style-minimax-h3

2. Images plus audio (no video)

If you have an image dataset and want your character to talk in the right voice, you only need audio clips and simple transcripts. In my Flonne set, each audio clip has a caption with the spoken line in quotes and the background sound described, like:

fl0nn3C, She says, "Whoosh." jazz music in background

- Lora: https://civitai.red/models/2968446/angel-college-cutie-minimax-h3

- Huggingface: https://huggingface.co/datasets/CitronLegacy/FlonneCollegeCutie_H3

3. Music

H3 can learn music, not just visuals. My Nagatoro dance lora is trained on only 3 short clips (about 5 seconds each) with speech and music.

- Lora: https://civitai.red/models/2980216/nagatoro-dance-gambare-gambare-senpai-h3

- Huggingface: https://huggingface.co/datasets/CitronLegacy/NagaDance_H3

Disclaimer 1: this lora is pretty janky! But its an interesting experiment that hopefully can be learned from.

Disclaimer 2: It would be probably be better to use one of the new music/audio base models for music but I've haven't learned about that yet.

4. A mix of images and videos

I think adding images to a video dataset helps the lora learn a character's appearance in fewer steps.

- Lora: https://civitai.red/models/2943697/teen-titans-style-dc-comics-minimax-h3?modelVersionId=3344472

- Huggingface: https://huggingface.co/datasets/CitronLegacy/StarfireTT_H3

- Lora: https://civitai.red/models/2946033/supergirl-or-kara-zor-el-dc-comics-minimax-h3

- Huggingface: https://huggingface.co/datasets/CitronLegacy/Supergirl_H3

Training settings

On CivitAI, I use the default settings and never train for fewer than 3000 steps.

Locally, I use https://github.com/shootthesound/Fizgig with its recommended defaults, and change these:

- Dim/Alpha: 32

- Resolution: 1024

- Epochs: I adjust this depending on what I'm training

Example datasets

These are datasets I think are good examples of what's possible. Each lora has a few small problems, so treat them as a starting point that gets you to about 80% quality, not as perfect references. All captions start with a trigger word, followed by plain-language description.

Flonne College Cutie H3

Contents: 13 images (my own fan art) plus 25 official Disgaea audio clips, each with a caption.

Link: https://huggingface.co/datasets/CitronLegacy/FlonneCollegeCutie_H3

Naga Dance H3

Contents: 3 videos (my own fan animation of Nagatoro, 768x1152, 30 fps, about 5 s each, with audio from YouTube), each with a caption.

Link: https://huggingface.co/datasets/CitronLegacy/NagaDance_H3

Starfire Teen Titans H3

Contents: 10 images and 8 video clips (1.6 to 6.2 s) from the official Teen Titans TV show, each with a caption.

Link: https://huggingface.co/datasets/CitronLegacy/StarfireTT_H3

Supergirl (Kara Zor-El) H3

Contents: 8 images (my own fan art, 832x1216) and 4 video clips (0.7 to 7.7 s, 1920x1080) from the Justice League Unlimited TV show, each with a caption.

Link: https://huggingface.co/datasets/CitronLegacy/Supergirl_H3

Quick Tip on writing prompts

I'm still not great at writing prompts for H3 so I go to ChatGPT or Grok (preferably) give it this link and tell it "Based on this guide give me a prompt for ________"
https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_base_en.md

Closing

I'd also like to highlight this post from koyan0418439: https://civitai.red/images/144182495. It's so cool that we can make things like this. Imagine your favorite character jamming to their theme song!

Images are 2D, but video feels like more than 3D because it has action and sound too!

PS

I'm hosting two Crucibles please check them out:

https://civitai.red/crucibles/18/pusheen-the-limits-of-ai

https://civitai.red/crucibles/17/ret-2-go-shantae-competition

Entry fee is at 10 buzz which is as low as it can go.

10