Download
1 variant available
This checkpoint includes a config file, download and place it along side the checkpoint.
Стартовый релиз оптимизированного воркфлоу для LTX-Video 2.5 / Initial release of the optimized LTX-Video 2.5 workflow.

50 1 2 3 4 5 6 7 8 9
200 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 9
540 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 9
LTX Video 2.5 and its derivatives, including LoRAs and fine-tunes, are licensed by Lightricks Ltd. under the LTX-2.x Community License Agreement and must be redistributed under that same agreement, with a copy included. Use is subject to the use restrictions in its Attachment A. Entities with annual revenues of at least $10,000,000 must obtain a paid commercial license from Lightricks before any commercial use.
Привет! Перед вами переработанный рабочий процесс для LTX-Video 2.5, который я специально заточил под максимальное качество изображения и стабильность кадра.
Воркфлоу собран на стандартных узлах и работает сразу «из коробки» — вам достаточно просто скачать модели и обновить недостающие ноды через встроенный поиск (Manager) вашей программы. Процесс не идеален (модель новая, капризная, со старой внутренней математикой латента и откровенно плохими энкодерами), поэтому здесь нет автопромптера — пишем всё руками. По этой причине очень рекомендую внимательно прочитать раздел про промпты и системные команды — без них модель начинает сильно халявить.
Мои личные впечатления от LTX-Video 2.5 🧠
Что понравилось:
• Дикая скорость работы и отличная оптимизация «из коробки» (модель просто летает) при низких разрешениях.
• Честная мультиязычность (отлично понимает русский язык без костылей-переводчиков).
• Прекрасная работа с нестандартными разрешениями и пропорциями картинок на входе — и это её главный, жирный плюс.
• Переработанная генерация звука: аудиопоток стал качественнее, чище и намного более предсказуем в генерации.
Что плохо:
• Откровенно «ленивая» и сломанная физика — модель постоянно срезает углы, избегает сложных движений и идет по самому легкому пути.
• Поведение звука при плохом промпте: несмотря на общее улучшение аудио-движка, если вы проигнорируете описание звука в промпте или опишете его слишком скудно, модель начинает сильно галлюцинировать и выдавать дикий белый шум. Промт для звука лучше всего писать вручную и максимально подробно.
• Стандартный рабочий процесс, который предлагают разработчики под эту модель, откровенно ущербен. Базовые ноды ужасно работают с латентным пространством: сначала они жестко роняют исходник в мелкий латент, крутят там кашу, а потом пытаются агрессивно натянуть это обратно встроенным апскейлером. На выходе получается сплошное мыло по краям и артефакты диффузии.
Моё решение проблемы 🛠️
Я полностью вырезал из генератора встроенный латентный апскейл. С учетом крутой оптимизации самой модели, этот шаг освободил тонну VRAM и позволил зайти с другой стороны — использовать сверхвысокие разрешения прямо на входе.
Воркфлоу построен на автоматической подготовке изображения: сначала входной кадр разгоняется в избыточные 2К/3К (через Swin2SR X2 и GLSL-шейдер контроля резкости). Модель сжимает уже сверхплотную картинку, и ей просто не хватает сил испортить детализацию в латенте до конца. На выходе видеопоток дополнительно обрабатывается оптимизатором изображения (шейдер резкости на 0 МБ нейросетевого VRAM + каскад RIFE), что дает на выходе честные, плотные ~1.5К премиального качества, где глаза, лица и контуры остаются целыми.
Важный момент: этот процесс лучше всего работает на роликах 5–10 секунд. Именно на такой длине модель выдаёт самое стабильное и красивое видео. Если делать длиннее (тестировал вплоть до 15-20 секунд) — качество прорисованных деталей (например, рук человека) постепенно проседает, появляются глюки и артефакты.
Главные фишки схемы 🔥
1. Избыточное 2K/3K наполнение на входе — обходим кривую латентную математику LTX 2.5 и заставляем её работать с честными пикселями.
2. Универсальность и гибкость — базово процесс настроен на видеокарты от 12 ГБ VRAM и выше, чтобы сразу выдавать максимальное 2К/3К качество. Однако схема способна стабильно и безопасно работать и на более слабых по мощности картах. Все ручные инструкции и подсказки по переключению проводов зашиты прямо внутри самого процесса в текстовых blocks — настройка займет пару минут.
3. Борьба с «пьяной» кособокой камерой — выведенная на опытах система команд, которая намертво прибивает камеру к полу и заставляет обленившуюся нейросеть отрабатывать физику объектов, а не тупо двигать холст на зрителя.
4. Плавные 48 FPS на выходе — финальный поток стабилизируется легким математическим шейдером резкости и нодой RIFE VFI.
Теперь о плохом (честно про железо) ⚠️
• Видеокарта: благодаря использованию квантованных int8 моделей (UNET и текстовый энкодер Gemma4-12b) и каскада Tiled VAE (плиточное декодирование 512), этот процесс стал всеядным. На 12 ГБ+ и выше он спокойно пережевывает 2К и 3К разрешения. Но даже на народных картах с 6–8 ГБ VRAM после быстрой ручной перестройки процесс не захлебнется и пойдет без ошибок Out of Memory (OOM) — разница будет только во времени рендера.
• Оперативная память (ВАЖНО): процесс безопасно работает с ФУЛЛ (максимальным) разрешением изображений только при 32 ГБ оперативной памяти и выше. Если вашей системной оперативной памяти меньше 32 ГБ, обязательно снизьте входное разрешение в нодах, иначе ComfyUI начнет сильно тормозить, улетит в своп или вылетит. С большим объемом ОЗУ всё работает спокойно и без нервов.
• Диск: строго желателен быстрый NVMe SSD от 512 ГБ (а лучше от терабайта), особенно учитывая тяжелые веса новых моделей.
......................................................
Hello! This is a reworked ComfyUI workflow for LTX-Video 2.5, which I specifically fine-tuned to achieve maximum image quality and frame stability.
The workflow is built entirely on standard nodes and works right "out of the box" — all you need to do is download the models and update any missing custom nodes using the built-in Manager search. The process is not perfect (the model is new, temperamental, runs on old internal latent math, and features noticeably poor text encoders), so there is no prompt expander here — we write everything manually. For this reason, I highly recommend carefully reading the section on prompts and system commands — without them, the model gets extremely lazy.
My Personal Impressions of LTX-Video 2.5 🧠
What I liked:
• Incredible generation speed and great optimization "out of the box" (the model absolutely flies) at lower resolutions.
• True multilingual support (it understands Russian perfectly without any translation node crutches).
• Excellent handling of custom resolutions and aspect ratios on the input image — and this is its biggest, most significant advantage.
• Reworked audio generation: the audio output is now higher quality, cleaner, and much more predictable during generation.
What I disliked:
• Pathetically "lazy" and broken physics — the model constantly cuts corners, avoids complex movements, and always takes the easiest path.
• Audio behavior with bad prompts: despite general improvements to the sound engine, if you ignore the audio description in your prompt or describe it poorly, the model starts hallucinating heavily and outputs loud white noise. It is best to write the audio prompt manually and as detailed as possible.
• The standard workflow provided by the developers for this model is downright terrible. The default nodes handle the latent space poorly: they heavily downscale the source image into a tiny latent, mess things up inside, and then try to aggressively stretch it back using the built-in upscale. This results in heavy blurring along the edges and terrible diffusion artifacts.
My Solution to the Problem 🛠️
I completely stripped the built-in latent upscale from the generator. Combined with the great optimization of the model itself, this step freed up a massive amount of VRAM and allowed me to approach it from the opposite side — using ultra-high resolutions directly at the input stage.
The workflow is built on automatic image preparation: first, the input frame is upscaled into an excessive 2K/3K resolution (using Swin2SR X2 and a custom GLSL shader for edge sharpness control). The model then compresses an already super-dense image, meaning it simply lacks the "strength" to destroy the details inside the latent space. At the final stage, the video stream is additionally processed by an image optimizer (a sharpness shader that consumes 0 MB of tensor VRAM + a RIFE interpolation cascade), giving you a clean, dense ~1.5K premium-quality output where eyes, faces, and object contours stay perfectly intact.
Important note: this workflow works best for short clips of 5–10 seconds. On this duration, the model delivers the most stable and beautiful video. If you make it longer (tested up to 15-20 seconds), the quality of fine details (like human hands) gradually degrades, and diffusion glitches/artifacts start to creep in.
Key Features of the Scheme 🔥
1. Excessive 2K/3K Input Stuffing — we bypass the broken latent math of LTX 2.5 and force it to work with true, dense pixels.
2. Universality and Flexibility — by default, the workflow is tuned for graphics cards with 12 GB VRAM and up to deliver maximum 2K/3K quality immediately. However, the scheme can safely and stably run on lower-spec cards as well. All manual instructions and guidelines on which wires to reconnect are built right into the workflow inside text notes — adjusting it takes just a couple of minutes.
3. Fighting the "Drunk" Camera — a system of commands discovered through extensive testing that nails the camera to the floor, forcing the lazy AI to actually process object physics instead of just lazily sliding the whole canvas toward the viewer.
4. Smooth 48 FPS Output — the final stream is stabilized via a lightweight mathematical sharpness shader and a RIFE VFI node.
Now for the Downsides (Honest Hardware Talk) ⚠️
• Graphics Card: thanks to the use of quantized int8 models (UNET and Gemma4-12b text encoder) alongside a Tiled VAE cascade (tile decode set to 512), this workflow is completely unpicky. On 12 GB+ cards and higher, it easily chews through 2K and 3K resolutions. But even on "budget" cards with 6–8 GB VRAM, after a quick manual adjustment, the process won't choke and will run completely free of Out of Memory (OOM) errors — the only difference will be the render time (tested on 3K while streaming YouTube and a TV show in the background — it holds rock-solid).
• RAM (CRITICAL): the workflow runs safely at FULL (maximum) resolution only if you have 32 GB of system RAM or higher. If your system RAM is less than 32 GB, make sure to lower the input resolution in the nodes, otherwise ComfyUI will lag heavily, hit the swap file, or crash. With 32 GB+ of RAM, everything runs smoothly and stress-free.
• Disk: a fast NVMe SSD of 512 GB (ideally 1 TB or more) is highly recommended, especially considering the heavy file sizes of the new models.