Sign In

Z-Image

Text-to-ImageBy Alibaba · Tongyi LabOpen weights

Z-Image is a family of open-weight text-to-image models from Alibaba's Tongyi Lab, built on a compact ~6B architecture that punches well above its size on prompt adherence, clean composition, and legible in-image text — including Chinese and English. The Turbo variant renders in just a handful of steps, while the Base model trades speed for maximum fidelity. Generate with both right here on Civitai — no GPU, no install.

7K+Z-Image models
79M+Images generated
6K+Z-Image LoRAs

About Z-Image

Z-Image is an open-weight text-to-image family from Alibaba's Tongyi Lab (released under the Tongyi-MAI banner), built on a compact ~6-billion-parameter architecture. Despite its small size it targets photorealistic image generation, bilingual text rendering in both English and Chinese, and robust instruction adherence — the kind of prompt-following that usually demands a much larger model. Because the weights are lightweight, the Turbo variant fits comfortably within 16GB of consumer VRAM and reaches sub-second latency on enterprise H800 GPUs.

The family ships as three checkpoints tuned for different jobs. Z-Image Turbo is a distilled model that produces a finished image in just 8 sampling steps (NFEs), making it the fast, few-step default; Z-Image Base is the non-distilled foundation model, released to unlock the full quality ceiling and to give the community a clean base for fine-tuning and custom development; and Z-Image Edit is a separate variant tuned for instruction-driven image-to-image editing. On Civitai, Turbo and Base are both hosted, so you can iterate quickly on Turbo and switch to Base when you want maximum fidelity — no downloads, no local GPU.

Reach for Z-Image when you want faithful prompt adherence and clean, legible in-image text — especially mixed English/Chinese text — at a fraction of the cost and wait of heavier models. Turbo is the natural pick for fast drafting and high-volume work, while Base rewards patience with the full step count for its best output. For the deepest style and character LoRA libraries, the SDXL-based Pony and Illustrious ecosystems still lead; Z-Image's own fine-tune library is smaller but growing, and its speed-to-quality ratio makes it a strong everyday text-to-image workhorse.

How to prompt Z-Image

  • Write in natural language and follow Z-Image's 6-part structure: Subject, Scene, Composition, Lighting, Style, Constraints — in that order. Lead with the subject (and any text you want rendered), since it matters most.
  • Keep prompts short: attention fades after roughly 75 tokens (about 50–60 words), so front-load the important content and trim trailing detail that will otherwise be ignored.
  • Do not use negative prompts. Turbo is a few-step distilled model with no CFG at inference, so every constraint has to be phrased positively inside the main prompt — describe what you want, not what you want to avoid.
  • Skip weight syntax like (word:1.3) — it is not supported. Control emphasis through word order and description instead.
  • For photorealism, lighting is the single strongest lever — be specific about it — and add sensory detail such as "skin texture," "fabric detail," "imperfections," or "film grain." For text in the image, Z-Image renders English and Chinese directly, so just state the words you want.

Featured Z-Image models

Curated — the models worth generating with first.

Z-Image Turbo — Z-Image checkpoint preview
Checkpoint
Z-Image Turbo

89K

3.0M

Civitai-hosted · fast few-step default

Z-Image Base — Z-Image checkpoint preview
Checkpoint
Z-Image Base

18K

153K

Civitai-hosted · max fidelity

Stickers.Redmond — Z-Image lora preview
StudioGhibli.Redmond — Z-Image lora preview
LuisaP Pixel Art Refiner — Z-Image lora preview
Soothing Atmosphere — Z-Image lora preview

Popular Z-Image LoRAs & add-ons

Top LoRAs by downloads — live data, refreshed daily. Stack them on any checkpoint.

Example generations

Curated, safe-for-work showcase — every image ships with its prompt and settings.

Z-Image example image: Gouache-painted portrait of a man with fluffy black hair leaning against a brick wall
Gouache-painted portrait of a man with fluffy black hair leaning against a brick wall
Z-Image Turbo · 832×1216
Z-Image example image: Photorealistic 35mm cinematic film still, shallow f/2.8 depth of field
Photorealistic 35mm cinematic film still, shallow f/2.8 depth of field
Z-Image Turbo · 832×1216
Z-Image example image: Realistic portrait, soft natural light, masterpiece-quality detail
Realistic portrait, soft natural light, masterpiece-quality detail
Z-Image Turbo · 832×1216
Z-Image example image: Avant-garde full-body fashion editorial, hyper-detailed modern couture
Avant-garde full-body fashion editorial, hyper-detailed modern couture
Z-Image Turbo · 832×1216
Z-Image example image: High-fashion editorial portrait, bold styling and dramatic studio lighting
High-fashion editorial portrait, bold styling and dramatic studio lighting
Z-Image Turbo · 832×1216
Z-Image example image: Gouache-and-ink cyborg woman on textured paper, half body of gears and circuitry
Gouache-and-ink cyborg woman on textured paper, half body of gears and circuitry
Z-Image Turbo · 832×1216

How to run Z-Image

Two paths — one takes ten seconds, one takes an afternoon.

🖥️ Run it locally

For power users who want full control.

Full control over the workflow
Batch and automate
Needs a 8GB+ VRAM GPU
Download ~12GB of weights
Set up ComfyUI yourself

No graphics card? The Civitai path above skips all of this.

Z-Image vs other ecosystems

Backed by Civitai usage data.

FeatureZ-ImageFluxSDXLQwen
Best forFast, faithful text-to-imagePhotorealism & textGeneral purpose, speedPrompt adherence & text
Prompt adherenceVery goodExcellentGoodExcellent
Text in imagesStrong (EN + CN)StrongWeakStrong
Speed on CivitaiFastest (few-step)Fast (4–8s)Fast (2–4s)Medium (6–12s)
LoRA ecosystem9,000+40,000+38,000+Growing
Available on Civitai✓ Yes✓ Yes✓ Yes✓ Yes

Frequently asked questions

How much does it cost to generate with Z-Image?

Generation on Civitai runs on Buzz, and Z-Image is one of the more affordable options — its compact ~6B architecture and few-step Turbo variant mean each image costs less Buzz than heavier models like Flux. Every account earns free Blue Buzz daily by reacting to images and other on-site activity, and because Z-Image is so light that daily Blue Buzz stretches a long way, letting you generate a lot without spending real money. Lean on Turbo for high-volume, few-step drafting; switch to Base when you want full fidelity and don't mind a slightly higher per-image cost, or add a membership for higher limits.

What's the difference between Z-Image Turbo and Z-Image Base?

Turbo is distilled for speed, producing images in just a handful of steps, while Base runs the full step count for maximum fidelity. Both are hosted on Civitai — try each and remix an example to compare.

Who made Z-Image?

Z-Image is developed and open-weighted by Alibaba's Tongyi Lab. You can run the official models on Civitai without downloading anything.

Can I train my own Z-Image LoRA?

Yes. Z-Image supports LoRA fine-tuning, and you can train one directly on Civitai — no local GPU needed. Publish it to earn Buzz when others generate with it.

Do I need a GPU to run Z-Image?

Not on Civitai — we run the compute for you. Locally, the ~6B weights fit comfortably on an 8GB+ VRAM GPU. Skip the setup and generate here instead.

Start generating with Z-Image now

No installation. No GPU. Runs in your browser.

Want more daily generations and a faster queue? Explore Civitai membership.