Download
1 variant available
This checkpoint includes a config file, download and place it along side the checkpoint.
Type
Workflows
Stats
2990 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 9
Reviews
(12)
Published
Sep 11, 2026
Base Model
MiniMax H3
Hash
AutoV2
F454E925B8

2150 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 9
1.1K0 1 2 3 4 5 6 7 8 9.0 1 2 3 4 5 6 7 8 9K
MiniMax H3 is licensed by MiniMax under the MiniMax H3 Community License Agreement. That agreement’s Applicable Territory excludes the European Union, the United Kingdom, the Republic of Korea and the United States of America. Your use of H3 and of any H3 derivative is subject to that agreement and its Acceptable Use Policy.
MiniMax H3
A MiniMax H3 workflow built on the fastvideo VSA distilled checkpoint, running
at 4 steps.
WHAT YOU CAN MAKE
Two generation nodes cover everything. Both ignore any input you leave
unplugged, so what you get depends on which assets you connect - not on which
branch you switch to.
T2V frame2v no assets
Image to Video frame2v first frame
First / Last Frame frame2v first frame + last frame
Reference to Video ref2v character images / reference video
Lip sync sits on top and works with either node.
ONE PANEL TO DRIVE IT
Every switch, every number and the prompt live in a single group: "0. Settings".
Your assets go in "1. Assets". Nothing else needs touching.
(1) Pick the model fast model, or int8 + Turbo LoRA
(2) Pick the generation node frame2v / ref2v
(3) Turbo LoRA on only with the int8 model
(4) How audio is used lip sync / guide / neither
(5) Which images are in use first frame, last frame, references
Size is set by aspect ratio x megapixels, with a cheat sheet in the graph.
Length is set in seconds - a math node rounds it to the 17n+5 frame counts H3
accepts, so you never have to count frames.
LIP SYNC
Give it a wav and the output audio is that file, bit for bit. The audio lane is
masked at 0, so it is never denoised - and because the model sees the finished
audio at every step while the picture is still noise, the mouth is drawn from
the sound rather than matched to it afterwards. It is inpainting, with the audio
lane as the known region.
Two characters in one shot can be driven separately. Feed one combined wav and
put the timings and the speakers in the prompt - including what the silent one
is doing:
[Shot 1] at 00:01 the girl on the left speaks: <d>...</d>
Meanwhile the girl on the right keeps her lips closed in a still, gentle smile.
at 00:04 the girl on the right speaks: <d>...</d>
The audio does not have to match the video length. Shorter audio simply leaves
the rest silent.
There is also a guide mode, which steers the generated audio toward your file
instead of pinning it. That one is frame2v only.fastvideo VSA の蒸留モデルを 4ステップで回す、MiniMax H3 のワークフローです。
なにが作れるか
生成ノードは2つだけです。どちらも挿していない素材は無視するので、
作るものは「どの枝に切り替えるか」ではなく「どの素材をつなぐか」で決まります。
テキストから frame2v 素材なし
画像から frame2v 最初のコマ
最初と最後のコマ frame2v 最初のコマ + 最後のコマ
参照から ref2v キャラ参照(画像)/ビデオ参照(動画)
リップシンクはこの上に乗り、どちらのノードでも使えます。
触るのは1か所だけ
スイッチも数値もプロンプトも、すべて「0. 設定」に集めてあります。
素材は「1. 素材」に入れてください。ほかは触る必要がありません。
(1) モデルを選ぶ ファストモデル / int8 + Turbo LoRA
(2) 生成ノードを選ぶ frame2v / ref2v
(3) Turbo LoRA int8 のときだけオン
(4) 音声の使い方 リップシンク / ガイド / 使わない
(5) 使う画像 最初のコマ・最後のコマ・参照画像
サイズは「比率 × メガピクセル」で決めます(早見表をグラフ内に同梱)。
長さは秒で入れれば、数式ノードが H3 の 17n+5 に丸めます。コマ数を数える必要はありません。
リップシンク
wav を渡すと、出力の音はそのファイルそのままです(ビット単位で一致)。
音声レーンをマスク0で固定するので、一切denoiseされません。
そしてモデルは、映像がまだノイズの段階から完成した音を聞いているので、
口は音に合わせにいくのではなく、音から導かれます。
仕組みとしてはインペイントで、音声レーンが「既知の領域」にあたります。
二人が同時に写っていても、話者ごとに口を動かし分けられます。
音声は1本にまとめて渡し、プロンプトに時刻と話者を書いてください。
黙っているほうの様子も書くのがコツです。
[Shot 1] at 00:01 the girl on the left speaks: <d>...</d>
Meanwhile the girl on the right keeps her lips closed in a still, gentle smile.
at 00:04 the girl on the right speaks: <d>...</d>
音声の尺を動画に合わせる必要はありません。短ければ、残りは無音になるだけです。
ガイドという別のやり方も入れてあります。音を固定するのではなく、
生成される音を渡したファイルに寄せるものです。こちらは frame2v 専用です。
