Wan
Wan is a family of open-weight video models from Alibaba that turn a text prompt or a still image into short, cinematic clips with fluid, realistic motion. It covers both text-to-video and image-to-video, and anchors the deepest LoRA and motion-effect ecosystem in open video generation. Generate a clip from a description in seconds — no GPU, no install. Run every Wan model right here on Civitai.
About Wan
Wan is a family of open-weight video models from Alibaba, spanning both text-to-video (T2V) and image-to-video (I2V). The line is built around Alibaba's Wan-VAE, an efficient video autoencoder, and Wan 2.1 was notable as the first open video model able to render both Chinese and English text inside a clip. It scales across sizes — from a lightweight 1.3B T2V model that fits in about 8GB of VRAM up to the flagship 14B models — so the same ecosystem covers everything from quick consumer-GPU drafts to high-fidelity cinematic generation.
The versions diverge in architecture and focus. Wan 2.2 introduces a Mixture-of-Experts (MoE) design that splits the denoising process across specialized expert models, enlarging capacity at the same compute cost, and adds a fast 5B TI2V variant (16×16×4 compression via Wan2.2-VAE) that generates 720p at 24fps on a card like a 4090; its A14B T2V and I2V models were trained on substantially more data with curated cinematic-aesthetic labels for lighting, composition, and color. Wan 2.5 is the hosted, audio-aware generation — synchronized voices, ambient sound, and music alongside cinematic 1080p output — while Wan 2.7 pushes temporal consistency and subject stability, reducing flicker, distortion, and identity drift across frames.
Choose Wan when you want open, controllable video with strong native image-to-video and the deepest LoRA and motion-effect ecosystem in open video generation. Reach for the 2.2 A14B models as the open-weight workhorses for T2V and I2V, the 5B variant when speed matters, and the hosted 2.5 / 2.7 releases when you want the best prompt adherence, motion realism, or audio-synced clips. Against closed APIs like Kling you trade some polish for open weights, stackable LoRAs, and no local install when you run it here.
How to prompt Wan
- Write in natural language, not tags — describe a cinematic scene in the order subject → action → setting → lighting → camera. Wan reads full descriptions, and weight syntax like (word:1.5) is ignored.
- Always include a camera direction; a missing one is the most common weak spot. Use Wan's vocabulary: "slow zoom in," "camera pans left," "dolly shot," "tracking shot," "aerial drone shot," or "static camera."
- Control motion explicitly with intensity cues like "gentle breeze," "slow-motion," or "rapid movement," and keep a single clip to one continuous action — a short shot can't hold several sequential events.
- Unlike many open image models, Wan supports negative prompts — a good default is "blurry, distorted, low quality, watermark, static, morphing, deformed hands" to suppress common video artifacts.
- For image-to-video, prompt only the motion you want added, not the static scene already in your image — describe what should move or change, plus the camera move and lighting/mood. For text-to-video, layer quality modifiers such as "cinematic lighting," "film grain," "HDR," or "shallow depth of field."
AI models move fast — new versions ship often, and a model’s capabilities or Buzz cost can change. For the latest, check the model’s own page before you generate.
Popular Wan LoRAs & add-ons
Top LoRAs by downloads — live data, refreshed daily. Stack them on any checkpoint.
Example videos
Curated, safe-for-work showcase — every clip ships with its prompt and settings.
How to run Wan
Two paths — one takes ten seconds, one takes an afternoon.
⚡ Run on Civitai (Recommended)
The fastest way to start.
🖥️ Run it locally
For power users who want full control.
No graphics card? The Civitai path above skips all of this.
Wan vs other ecosystems
Backed by Civitai usage data.
| Feature | Wan | Hunyuan Video | LTXV | Kling |
|---|---|---|---|---|
| Best for | T2V + I2V, huge LoRA ecosystem | Cinematic text-to-video | Fast, lightweight clips | Polished cinematic clips (API) |
| Prompt adherence | Very good | Very good | Good | Excellent |
| Image-to-video | Native, strong | Limited | Yes | Yes |
| Speed on Civitai | Medium | Medium | Fast | Medium (API) |
| LoRA ecosystem | 2,500+ (largest for video) | Growing | Small | None (closed) |
| Available on Civitai | ✓ Yes | ✓ Yes | ✓ Yes | ✓ Yes |
Frequently asked questions
How much does it cost to generate with Wan?
What's the difference between Wan text-to-video and image-to-video?
Which Wan version should I use?
Can I use LoRAs and motion effects with Wan?
Do I need a GPU to run Wan?
Start generating with Wan now
No installation. No GPU. Runs in your browser.