Sign In

[Workflow] InfiniteTalk · Image to Video For AD Film

Updated: Aug 3, 2026

styleaifloyocomfyuii2vanimate face

Download

1 variant available

Config Other

InfiniteTalk · Image to Video For AD Film.json

126.91 KB

Verified:

Type
Workflows
Stats

17

Reviews

No reviews yet

Published

Aug 3, 2026

Base Model

Wan Video 14B i2v 720p

Hash
AutoV2
747A9CDAB9
default creator card background decoration
Followers - 492

492

Likes - 2347

2.3K

Holiday 2023: Golden Donor

License:

Apache 2.0
z-image_00054_.png

Turn a portrait and any-length audio into a talking video

Open this ready-to-run workflow on Floyo. No install needed.

HOW IT WORKS

Step 1. Upload your portrait. A front-facing headshot of the person you want to animate, well-lit, with the mouth and jaw visible. Works great with: portraits · headshots · AI-generated faces · character renders

Step 2. Upload your audio. Speech, narration, or singing, at any length. The workflow separates the vocals from background noise on its own before it starts. Works great with: voiceovers · podcast clips · dialogue · singing

Step 3. Write a short prompt. Describe the performance and mood, like "a man talking calmly with slight head nods." This steers the motion around the lip-sync. Keep it about the delivery, not the room.

Step 4. Hit run and download. InfiniteTalk generates the video in overlapping windows and tiles them across the whole track, then combines them into one MP4 at 25fps with your audio embedded. Ready for: Premiere · DaVinci Resolve · After Effects · YouTube · TikTok · podcasts

First time? Upload a front-facing portrait and an audio clip, leave every setting as-is, and hit run. The defaults handle the rest.

Overview

This workflow turns a portrait and an audio file into a lip-synced talking video using Wan 2.1 14B with the InfiniteTalk adapter. The point of it is length. Most lip-sync tools cap out at a few seconds, because they generate a fixed number of frames in one pass. InfiniteTalk processes your audio in overlapping 81-frame windows and loops until the full recording is covered, so a 10-second clip and a 5-minute monologue run through the same pipeline with no duration limit. It separates vocals from background music, reads the speech with a Wav2Vec2 encoder, and drives the mouth from the sound, so the sync works across languages. You upload a portrait and audio, write a short prompt, and get the full-length video back. Generation time scales with audio length, and a short clip takes about 2 minutes 22 seconds. No setup, no nodes to wire.

Who it's for: creators, podcasters, and course makers who want a talking avatar for long audio in ComfyUI without wiring the windowed lip-sync pipeline from scratch. Not for: a fast one-off clip or a profile-angle portrait. It is a heavy render, and the mouth and jaw need to be visible and front-facing.

Why Floyo

Floyo is the only ComfyUI platform built for teams in the browser.

  • Made for teams. Share run history, files, and models across your whole team. A teammate opens your exact run and picks up where you left off. No file handoffs, no version confusion.

  • No install, no setup. Every workflow and model is preloaded. Open it in your browser and run. Nothing to download, nothing to configure.

  • No local hardware. Workflows run on H100 NVL GPUs, so heavy models run fast without a card of your own. Your VRAM stops being the limit.

  • Open and closed models in one place. Floyo runs open-source workflows and API models side by side.

How to use (in your browser on Floyo)

  1. Upload your front-facing portrait and an audio file of any length.

  2. Write a short prompt describing the performance, then run.

  3. Download the talking video with your audio embedded, ready for any editor.

Expectations The models are preloaded, so there is nothing to download. Running it needs a free Floyo account. This is a heavy render, and generation time scales with your audio: a short clip is about 2 minutes 22 seconds, and longer tracks take proportionally longer. A few things decide the result. Front-facing portraits with a clear, unobstructed mouth sync best, while profile angles and hair or hands over the lower face do not. Clean, isolated vocals give tighter sync than audio buried under music or reverb, even though the workflow separates vocals first. To push the mouth harder, raise the audio_scale value. On licensing: InfiniteTalk builds on Alibaba's open-source Wan 2.1, but the adapter, LoRAs, and supporting models each carry their own terms, so check each component before commercial use.

Use Cases

Podcasts & Narration. Turn a full episode or a long narration into a talking avatar video, giving audio-only content a face for YouTube and social.

Course & Explainer Content. Generate a consistent instructor from one portrait across full-length lessons, no filming.

Singing & Music. Feed a vocal track and a character image to produce a singing performance synced for the whole song.

Long-Form Social. Make talking clips longer than other lip-sync tools allow, covering full walkthroughs, story times, and monologues.

What can InfiniteTalk turn into a talking video?