Sign In

Wan

Text-to-Video & Image-to-VideoBy Alibaba (Wan)Open weights

Wan is a family of open-weight video models from Alibaba that turn a text prompt or a still image into short, cinematic clips with fluid, realistic motion. It covers both text-to-video and image-to-video, and anchors the deepest LoRA and motion-effect ecosystem in open video generation. Generate a clip from a description in seconds — no GPU, no install. Run every Wan model right here on Civitai.

3K+Wan models
11M+Videos generated
1K+Wan LoRAs

About Wan

Wan is a family of open-weight video models from Alibaba, spanning both text-to-video (T2V) and image-to-video (I2V). The line is built around Alibaba's Wan-VAE, an efficient video autoencoder, and Wan 2.1 was notable as the first open video model able to render both Chinese and English text inside a clip. It scales across sizes — from a lightweight 1.3B T2V model that fits in about 8GB of VRAM up to the flagship 14B models — so the same ecosystem covers everything from quick consumer-GPU drafts to high-fidelity cinematic generation.

The versions diverge in architecture and focus. Wan 2.2 introduces a Mixture-of-Experts (MoE) design that splits the denoising process across specialized expert models, enlarging capacity at the same compute cost, and adds a fast 5B TI2V variant (16×16×4 compression via Wan2.2-VAE) that generates 720p at 24fps on a card like a 4090; its A14B T2V and I2V models were trained on substantially more data with curated cinematic-aesthetic labels for lighting, composition, and color. Wan 2.5 is the hosted, audio-aware generation — synchronized voices, ambient sound, and music alongside cinematic 1080p output — while Wan 2.7 pushes temporal consistency and subject stability, reducing flicker, distortion, and identity drift across frames.

Choose Wan when you want open, controllable video with strong native image-to-video and the deepest LoRA and motion-effect ecosystem in open video generation. Reach for the 2.2 A14B models as the open-weight workhorses for T2V and I2V, the 5B variant when speed matters, and the hosted 2.5 / 2.7 releases when you want the best prompt adherence, motion realism, or audio-synced clips. Against closed APIs like Kling you trade some polish for open weights, stackable LoRAs, and no local install when you run it here.

How to prompt Wan

  • Write in natural language, not tags — describe a cinematic scene in the order subject → action → setting → lighting → camera. Wan reads full descriptions, and weight syntax like (word:1.5) is ignored.
  • Always include a camera direction; a missing one is the most common weak spot. Use Wan's vocabulary: "slow zoom in," "camera pans left," "dolly shot," "tracking shot," "aerial drone shot," or "static camera."
  • Control motion explicitly with intensity cues like "gentle breeze," "slow-motion," or "rapid movement," and keep a single clip to one continuous action — a short shot can't hold several sequential events.
  • Unlike many open image models, Wan supports negative prompts — a good default is "blurry, distorted, low quality, watermark, static, morphing, deformed hands" to suppress common video artifacts.
  • For image-to-video, prompt only the motion you want added, not the static scene already in your image — describe what should move or change, plus the camera move and lighting/mood. For text-to-video, layer quality modifiers such as "cinematic lighting," "film grain," "HDR," or "shallow depth of field."

AI models move fast — new versions ship often, and a model’s capabilities or Buzz cost can change. For the latest, check the model’s own page before you generate.

Featured Wan models

Curated — the models worth generating with first.

Checkpoint
Wan 2.2 T2V-A14B

25K

2K

Civitai-hosted · default

Checkpoint
Wan 2.5 T2V

565

Civitai-hosted · latest

Checkpoint
Wan 2.7

5K

Civitai-hosted · newest

Checkpoint
Wan 2.1 14B T2V

27K

1.6M

Civitai-hosted · open weights

Popular Wan LoRAs & add-ons

Top LoRAs by downloads — live data, refreshed daily. Stack them on any checkpoint.

Example videos

Curated, safe-for-work showcase — every clip ships with its prompt and settings.

Cinematic nighttime chase through a rain-soaked futuristic street market
Wan 2.2 T2V · 480×832 · 16fps
Six dancers in a synchronized rooftop routine over a coastal city at sunset
Wan 2.2 T2V · 480×832 · 16fps
A tiny frog hopping up and down a large leaf, wet from rain
Wan 2.2 T2V · 480×832 · 16fps
A gothic vampire portrait with subtle, moody motion
Wan 2.2 I2V · 480×832 · 16fps
A magical-girl heroine in a bright transformation-style scene
Wan 2.2 I2V · 480×832 · 16fps
An elegant character gliding toward the camera on a sunlit seaside balcony
Wan 2.2 I2V · 480×832 · 16fps

How to run Wan

Two paths — one takes ten seconds, one takes an afternoon.

🖥️ Run it locally

For power users who want full control.

Full control over the workflow
Batch and automate
Needs a 16GB+ VRAM (14B); less for 1.3B / 5B GPU
Download ~16–32GB (14B) of weights
Set up ComfyUI yourself

No graphics card? The Civitai path above skips all of this.

Wan vs other ecosystems

Backed by Civitai usage data.

FeatureWanHunyuan VideoLTXVKling
Best forT2V + I2V, huge LoRA ecosystemCinematic text-to-videoFast, lightweight clipsPolished cinematic clips (API)
Prompt adherenceVery goodVery goodGoodExcellent
Image-to-videoNative, strongLimitedYesYes
Speed on CivitaiMediumMediumFastMedium (API)
LoRA ecosystem2,500+ (largest for video)GrowingSmallNone (closed)
Available on Civitai✓ Yes✓ Yes✓ Yes✓ Yes

Frequently asked questions

How much does it cost to generate with Wan?

Generation on Civitai runs on Buzz, and you can claim free Blue Buzz every day — through actions like reacting to images and other on-site activity — to put straight toward generating, no real money required. Because Wan spans a wide size range, cost tracks the model: the lighter 1.3B and 5B variants are cheap per clip so your daily Blue Buzz stretches far, while the 14B and newer hosted 2.5 / 2.7 releases cost more Buzz per generation, so heavier use means letting your Blue Buzz accumulate or adding a membership for higher limits.

What's the difference between Wan text-to-video and image-to-video?

Text-to-video (T2V) builds a clip from a written prompt, while image-to-video (I2V) animates a still image you provide. Wan does both, and you can pick either mode right in the Civitai generator.

Which Wan version should I use?

Wan 2.2 is the open-weight workhorse for T2V and I2V; 2.5 and 2.7 are the newer hosted releases with improved motion and detail. All of them are generatable on Civitai — try a prompt on each and compare.

Can I use LoRAs and motion effects with Wan?

Yes. Civitai hosts 2,500+ Wan LoRAs — including motion effects like 360° rotation and squish — that you can stack in the generator. Remix any example above to see how the settings carry over.

Do I need a GPU to run Wan?

Not on Civitai — we run the compute for you. Locally the 14B models want a 16GB+ VRAM GPU in ComfyUI, though the lighter 1.3B and 5B variants run on less. Skip the setup and generate in the browser.

Start generating with Wan now

No installation. No GPU. Runs in your browser.

Want more daily generations and a faster queue? Explore Civitai membership.