Sign In

# LTX-2.3 American Accent LoRA — make accent prompts actually work

98

Updated: Aug 4, 2026

base model

Download

1 variant available

bf16 SafeTensor

ltx23_accent_american_v2_rank32.safetensors

BF16, good balance • 168.35 MB

Verified:

Type
LoRA
Stats

420

98

363

Reviews
Published

Jul 30, 2026

Base Model

LTXV 2.3

Hash
AutoV2
1AEC71F318
default creator card background decoration
Followers - 128

128

Likes - 410

410

Bronze Base model Badge

Everything here is free and stays free — the format spec, the nodes, the workflows, the cartridges, the LoRAs. If it saved you a night of debugging (it contains several hundred of mine), tips keep the 5090 warm:

LTX-2.3 American Accent LoRA — make accent prompts actually work

Makes accent wording in your prompt actually work. LTX-2.3's voice prior ignores accent requests in exactly the regions where you need them most — young female characters default to Australian-leaning voices even when the prompt says "in a casual American accent". This LoRA turns that wording into a reliable control.

Read this first: the 24 fps rule

This LoRA cannot help you at the wrong frame rate. LTX-2.3's joint audio-video prior is 24 fps-native, and render fps is a hidden accent dial: at 25 fps the same prompt and seed render non-rhotic southern British, at 30 fps broad Australian — and off-24 fps overrides accent wording entirely, LoRA or no LoRA. Set your workflow's fps/frame_rate widgets to 24, then use this LoRA. (Dose-response verified by A/B on identical configs.)

What it does — measured

A/B matrix on the hardest region (young pale-skinned woman, casual camcorder monologue), 24 fps, two seeds per cell, blind phonetic review:

spoken-line wordingwithout this LoRAwith it @ 1.0nothing statedAustralianAustralian"saying in a casual American accent"Australian (both seeds)General Americanrich scaffold (voice timbre + accent binding)Australian (both seeds)General American

Read the top row carefully: this LoRA does not force American unconditionally. It makes the model obey the accent you ask for where the base model refuses. Ask for nothing and you still get the base model's lean — so state the accent on every spoken line and let the LoRA do the enforcing.

Known ceiling — read before you try for a regional accent

It enforces General American and flattens regional sub-flavours. "Soft Southern" and "Boston-flavored" both render as General American in testing. If you need a specific regional dialect, this is not that tool yet — that needs dialect-labelled training data, not different prompt wording.

Lip-sync safe by construction

1,152 LoRA tensors, all in the audio branches (audio_attn1 / audio_attn2 / audio_ff). Zero video tensors, zero cross-modal tensors. Faces, video content, and lip motion are mathematically untouched.

Usage

  • Strength 1.0 in any LTX-2.3 LoRA loader. Verified in long multishot production runs alongside a video-branch LoRA with no interference.

  • Compatible with the LTX-2.3 22B family: stock distilled 1.1 and the JoyAI-Echo surgical merges.

  • Render at 24 fps. Say the accent. That is the whole recipe.

Training

ai-toolkit, rank 32 / alpha 32, 3,000 steps, lr 1e-4, qfloat8. American-English read speech from LibriSpeech (CC BY 4.0), muxed over static video so only the audio lane carries signal; captions are verbatim transcripts with the accent deliberately unnamed, so the American prior trains as always-on behavior rather than a trigger phrase. The audio-branch-only module scope keeps the static training video harmless — the video branches never receive a gradient.

Also on HuggingFace: joeygambino/ltx23-accent-american-audio-lora