Sign In

MiniMax H3 Reference to Video - keep the same person in any scene

Download

1 variant available

Config Other

stableyogi-h3-reference.json

28.99 KB

Verified:

Type
Workflows
Stats

417

Reviews
Published

Aug 28, 2026

Base Model

MiniMax H3

Hash
AutoV2
1F1C7108DD
default creator card background decoration
Reactions - 1056510

1.1M

Followers - 31383

31.4K

Likes - 163467

163.5K

Gold β€œSSS-Rank” Command Crest

MiniMax H3 is licensed by MiniMax under the MiniMax H3 Community License Agreement. That agreement’s Applicable Territory excludes the European Union, the United Kingdom, the Republic of Korea and the United States of America. Your use of H3 and of any H3 derivative is subject to that agreement and its Acceptable Use Policy.

MiniMax H3

Hand it one photo of a person. It gives you back a completely different shot, in a place you never photographed, with that same person in it. Your photo never shows up in the result. It is read, and then set aside.

That is not the same thing as image to video, even though almost everyone starts by assuming it is. Image to video takes your picture and makes it move. This reads your picture, puts it down, and writes a new scene from scratch with that face in it.

🎬 Watch it built step by step: the full walkthrough on YouTube. Every setting, in order, on a normal home machine.

Two files in this post

  • Standard build - 20 steps. The one to start with if you want the safest result.

  • Turbo build - 8 steps, and one extra file to download. On my machine a 10 second clip took 153 seconds here against 349 seconds on the standard build. Same person, same picture quality. I ran both on the same seed before putting this up.

What you need

The folder is the part people get wrong. A correct file in the wrong folder looks exactly like a missing model, and ComfyUI will not tell you which.

Two things worth knowing

  • The prompt does far more work here than in the other H3 workflows. You are describing a whole scene, not a movement, because nothing about the scene comes from your photo.

  • Both decoders are needed. One is for the picture and one is for the sound, and leaving either out fails in a way that does not obviously say so.

I wrote up every setting, what each one does and the mistakes that cost me time, in the full written guide. It is free and there is no card.

There is also a paid course that goes further with this model if you want it. Either way the workflows here are the whole thing, not a trimmed version.