Download
1 variant available
**v1.2 — A/V sync fix, AutoFinish, encoder picker, ready-to-run workflow**
This listing goes from v1.0 straight to v1.2. v1.1 was the workflow release, which now has its own listing.
**Breaking:** JoyEcho Model Loader's widgets were reorganized — all inputs are now optional, with file pickers above their manual-path fallbacks. Saved graphs from v1.0 will load this node with shifted values: delete and re-add the node (one output wire to reconnect), or start from the bundled workflow.
**Fixed**
- **Multi-shot A/V sync drift.** Each shot's generated audio runs ~30 ms shorter than its video, and the final assembly compounded that — audio ran ~120 ms early by shot 5, so lip sync visibly degraded through the piece. Each shot's wav is now padded to its own video duration before concat; sync is exact by construction. If multi-shot renders drifted out of sync late in a piece, this was why.
- **LLM Enhance no longer demands an API key for local endpoints.** When base_url is a local address (localhost, 127.0.0.1, 192.168.x, .local), a placeholder is sent automatically — Ollama, LM Studio and llama.cpp ignore it anyway. Cloud providers still raise a clear error naming the endpoint. An empty field previously errored even against a local server, which was the most common first-run stumble.
- **Reference Picker loads scripted refs for non-character scenes.** A scene's JSON `refs` block (`{"kelpie": "shot.png"}`) is now honored directly, so creature and location references load without the folder name having to appear in the prompt prose. Previously only names written in the prose matched.
- Hires refine no longer crashes on resolution retarget (cached seqlen mismatch).
**New**
- **AutoFinish node + worker** (missing from v1.0 entirely): after the render it automatically upscales the per-shot masters and assembles the audio-synced final MP4 — no manual upscale/concat step.
- **Model Loader: `gemma_file` dropdown** — pick the text encoder from `models/text_encoders` or `models/clip` instead of typing a path. Single-file .safetensors or .gguf encoders supported (v1.0 required the HF directory).
- **Prompt Source: `manual_path`** — point at any .txt/.json prompt file anywhere on disk, no folder conventions needed.
- **Conditioning cache** — the slow Gemma encode is cached per (prompt, encoder); re-renders with unchanged text skip the encode entirely.
- **Reference Picker:** character dropdown that survives node refreshes; default refs folder is now `input/joyecho_refs/` (one subfolder per character).
- **Writer system prompts** gained two hard-won sections: FRAMING (use shot-type nouns — "framed from the waist up" gets ignored or mis-read; keep speaking shots at medium close-up or tighter, because the VAE packs 32 px into one latent token and a wide shot puts the mouth below one token) and HUMAN MOVEMENT (state gait/turn/contact mechanics; only describe body parts that are in frame). The short-story prompt file is now included alongside the long-story one.
Canonical source and updates: https://huggingface.co/joeygambino/joyai-echo-multishot-workflow
Show more

2070 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 9
6440 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 9
20.3K0 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 9.0 1 2 3 4 5 6 7 8 9K

Everything here is free and stays free — the format spec, the nodes, the workflows, the cartridges, the LoRAs. If it saved you a night of debugging (it contains several hundred of mine), tips keep the 5090 warm:
🔁 Liberapay (recurring)
⚡ Or right here: the Civitai tip button on this page sends Buzz directly.
🚨 v2.0 IS HERE — RIFTCAST: CHARACTERS ARE FILES NOW. Design a human from dropdowns, watch them audition, and get a portable character file anyone can reuse — zero training. Plus the 24fps accent discovery that fixes every "why is she suddenly British" bug. Full story in the v2.0 version notes. 🚨
JoyAI-Echo Multishot — the character video studio for LTX-2.3
One character. Any number of shots. Same face, same voice, every time — and as of 2.0, your characters are files you can share.
This is a complete local pipeline for character-driven video on LTX-2.3's joint audio-video model: it generates the picture AND the voice in one diffusion pass — no TTS chain, no lip-sync post, no per-word API costs. A cross-shot memory bank keeps identity and voice locked across an entire multi-shot production, and a set of hard-won fixes makes the base model behave in ways stock workflows can't.
RiftCast — characters are files now
Things come through the Rift. Now characters do too.
A .riftcast cartridge is one file carrying a character's voice (a 4-second anchor clip), face (reference stills), canonical description, and optionally their LoRAs and home environments. Drop it in input/riftcast/, restart, and mention the character in any script: they render with their face and their voice, in any scene, zero training. Cartridges extend the roleplay world's Character Card V3/CHARX lineage into photoreal video — a cartridge can carry a chat persona too, so the same character works in your roleplay client.
Three ways to get one:
Download —
WREN.riftcastships in this package; more on the HuggingFace repo.Cut one from any render you like — one command:
python riftcast.py cut <master.mp4> <NAME> <dna.txt>.Design one from scratch — the bundled RiftCast Studio workflow is a full character creator. A Character Designer node covers identity (gender, age, ethnicity, skin, height, build, voice timbre, accent) and a Style + Wardrobe node covers appearance: 71 styles across 10 families, plus hair colour, hair shape, makeup, accessories, demeanor and a freeform wardrobe override. Queue it and the character records an audition tape; the Packer cuts their anchor and reference stills from that render and installs the cartridge automatically. Don't like who showed up? Re-queue with a new seed. Like them? They're a file, forever. (Our test reviewer rated a Designer character's lip sync "Real — very high confidence.")
One workflow, one switch: RiftCast Studio routes between your prompt files (LPFF/JSON batch rendering, the classic path) and the Character Designer with a single dropdown.
The style dropdown never sends its own name
This is the part that makes it work rather than being a word list. "Goth" and "preppy" mean nothing to a video model, so no style label is ever written into a prompt. Each of the 71 entries expands into concrete renderable descriptors — garments named with material and condition, hair shape, makeup with placement, worn objects, and a demeanor that drives how the character physically carries themselves on camera. Picking one goth entry produces, in part:
"...hair backcombed high at the crown with a straight fringe cut level with the eyebrows, wearing a long-sleeved black velvet dress with a frayed hem over laddered fishnet tights, and buckled boots scuffed grey at the toe, matte pale foundation, black liner drawn thick and winged past the outer corner..."
Hair colour stays its own dropdown and styles specify only shape, so the two can never collide. Every style carries both a masculine and a feminine wardrobe reading, so it works across presentations instead of being gender-locked, and anything you set explicitly overrides the style's contribution.
The 24 fps law — why your accents broke
The single most important thing this package knows: LTX-2.3's joint audio-video prior is 24 fps-native, and render fps is a hidden accent dial. At 25 fps the same prompt and seed render non-rhotic southern British; at 30 fps, broad Australian — and off-24 fps overrides accent wording in your prompt entirely. Verified by A/B with blind phonetic review. Everything here defaults to 24 and warns when you stray. If you want a British or Australian character, render their scenes at 25/30 — it beats any wording. (Pairs with the American-accent audio LoRA, which makes accent wording enforceable in the young-voice registers the base model ignores.)
The rest of the studio
Cross-shot memory bank — identity and voice persist across shots; anchor+latest policy stops drift snowballs.
Voice casting — drop a clip in
joyecho_voices/<tag>/and that character speaks with that voice in every render. Script-pinned voices viavoice_refs.Finishing — AutoFinish builds your master automatically; deterministic upscale; optional
temporal_upscaledoubles motion to ~48 fps masters (24 fps render law preserved, audio untouched).Long takes — up to 1441 frames (60 s) single-shot; the old ~10 s lip-sync cliff is fixed at the RoPE-clock level (Bug fix #0).
Correctness — a full sampler-path audit (seeded hires, cache keyed on checkpoint, cloned banks), fp8/INT8/GGUF loading paths, and widget values that survive updates (saved by name, not position).
Hardware
Built and tested on RTX 5090/3090. GGUF DiT + the 9 GB VAE companion runs the whole stack in ~11 GB system RAM instead of ~60.
Everything is also on GitHub and HuggingFace — node pack, format spec, LoRAs (surface realism too), quants, and demo cartridges. LTX-2 Community License.
