Sign In

MiniMax H3 Fast - 4-step workflow with VSA and lip sync

Updated: Sep 15, 2026

toolminimaxh3workflow

Download

1 variant available

Config Other

MiniMaxH3_all_in_one_en_V2.json

159.37 KB

Verified:

Type
Workflows
Stats

299

Reviews
Published

Sep 11, 2026

Base Model

MiniMax H3

Hash
AutoV2
F454E925B8
default creator card background decoration
Followers - 215

215

Likes - 1073

1.1K

MiniMax H3 is licensed by MiniMax under the MiniMax H3 Community License Agreement. That agreement’s Applicable Territory excludes the European Union, the United Kingdom, the Republic of Korea and the United States of America. Your use of H3 and of any H3 derivative is subject to that agreement and its Acceptable Use Policy.

MiniMax H3

civitai.png
A MiniMax H3 workflow built on the fastvideo VSA distilled checkpoint, running
at 4 steps.

WHAT YOU CAN MAKE

Two generation nodes cover everything. Both ignore any input you leave
unplugged, so what you get depends on which assets you connect - not on which
branch you switch to.

  T2V                  frame2v    no assets
  Image to Video       frame2v    first frame
  First / Last Frame   frame2v    first frame + last frame
  Reference to Video   ref2v      character images / reference video

Lip sync sits on top and works with either node.


ONE PANEL TO DRIVE IT

Every switch, every number and the prompt live in a single group: "0. Settings".
Your assets go in "1. Assets". Nothing else needs touching.

  (1) Pick the model            fast model, or int8 + Turbo LoRA
  (2) Pick the generation node  frame2v / ref2v
  (3) Turbo LoRA                on only with the int8 model
  (4) How audio is used         lip sync / guide / neither
  (5) Which images are in use   first frame, last frame, references

Size is set by aspect ratio x megapixels, with a cheat sheet in the graph.
Length is set in seconds - a math node rounds it to the 17n+5 frame counts H3
accepts, so you never have to count frames.


LIP SYNC

Give it a wav and the output audio is that file, bit for bit. The audio lane is
masked at 0, so it is never denoised - and because the model sees the finished
audio at every step while the picture is still noise, the mouth is drawn from
the sound rather than matched to it afterwards. It is inpainting, with the audio
lane as the known region.

Two characters in one shot can be driven separately. Feed one combined wav and
put the timings and the speakers in the prompt - including what the silent one
is doing:

  [Shot 1] at 00:01 the girl on the left speaks: <d>...</d>
  Meanwhile the girl on the right keeps her lips closed in a still, gentle smile.
  at 00:04 the girl on the right speaks: <d>...</d>

The audio does not have to match the video length. Shorter audio simply leaves
the rest silent.

There is also a guide mode, which steers the generated audio toward your file
instead of pinning it. That one is frame2v only.
fastvideo VSA の蒸留モデルを 4ステップで回す、MiniMax H3 のワークフローです。

なにが作れるか

生成ノードは2つだけです。どちらも挿していない素材は無視するので、
作るものは「どの枝に切り替えるか」ではなく「どの素材をつなぐか」で決まります。

  テキストから          frame2v    素材なし
  画像から              frame2v    最初のコマ
  最初と最後のコマ      frame2v    最初のコマ + 最後のコマ
  参照から              ref2v      キャラ参照(画像)/ビデオ参照(動画)

リップシンクはこの上に乗り、どちらのノードでも使えます。


触るのは1か所だけ

スイッチも数値もプロンプトも、すべて「0. 設定」に集めてあります。
素材は「1. 素材」に入れてください。ほかは触る必要がありません。

  (1) モデルを選ぶ        ファストモデル / int8 + Turbo LoRA
  (2) 生成ノードを選ぶ    frame2v / ref2v
  (3) Turbo LoRA          int8 のときだけオン
  (4) 音声の使い方        リップシンク / ガイド / 使わない
  (5) 使う画像            最初のコマ・最後のコマ・参照画像

サイズは「比率 × メガピクセル」で決めます(早見表をグラフ内に同梱)。
長さは秒で入れれば、数式ノードが H3 の 17n+5 に丸めます。コマ数を数える必要はありません。


リップシンク

wav を渡すと、出力の音はそのファイルそのままです(ビット単位で一致)。
音声レーンをマスク0で固定するので、一切denoiseされません。
そしてモデルは、映像がまだノイズの段階から完成した音を聞いているので、
口は音に合わせにいくのではなく、音から導かれます。
仕組みとしてはインペイントで、音声レーンが「既知の領域」にあたります。

二人が同時に写っていても、話者ごとに口を動かし分けられます。
音声は1本にまとめて渡し、プロンプトに時刻と話者を書いてください。
黙っているほうの様子も書くのがコツです。

  [Shot 1] at 00:01 the girl on the left speaks: <d>...</d>
  Meanwhile the girl on the right keeps her lips closed in a still, gentle smile.
  at 00:04 the girl on the right speaks: <d>...</d>

音声の尺を動画に合わせる必要はありません。短ければ、残りは無音になるだけです。

ガイドという別のやり方も入れてあります。音を固定するのではなく、
生成される音を渡したファイルに寄せるものです。こちらは frame2v 専用です。