Generation Ecosystems
Every major AI image and video model family, hosted and generatable on Civitai. Pick an ecosystem to see its best checkpoints and LoRAs, real example generations with their prompts, and how it compares — then start creating right here.
Image models
FluxImage
Flux is a new generation of open-weight text-to-image models with class-leading prompt adherence, legible in-image text, and striking photorealism.Flux.2Image
Flux.2 is the latest generation of Black Forest Labs text-to-image models — sharper prompt adherence, cleaner in-image text, and stronger photorealism than the original Flux.SDXLImage
SDXL (Stable Diffusion XL) is Stability AI's open-weight text-to-image model and the backbone of the largest fine-tune and LoRA ecosystem in open image generation.PonyImage
Pony Diffusion V6 XL is an SDXL fine-tune by AstraliteHeart built for characters, anime, and stylized art with strong control over pose and composition through its tag-based "score_" prompt system.IllustriousImage
Illustrious is an SDXL-based text-to-image model built by OnomaAI for anime and illustration.NoobAIImage
NoobAI-XL (NAI-XL) is an Illustrious-based anime checkpoint trained by Laxhar Lab on a massive Danbooru/e621 dataset, giving it deep character knowledge and strong booru-tag prompt understanding.Nano BananaImage
Nano Banana is Google's Gemini-based image generation and editing model, best known for keeping a subject looking like themselves across edits — change an outfit, hairstyle, or setting while the face stays consistent.Imagen 4Image
Imagen 4 is Google DeepMind's latest text-to-image model, built to turn natural-language prompts into ultra-high-quality, photorealistic images with a focus on visual fidelity, fine detail, and compositional accuracy.SeedreamImage
Seedream is ByteDance's text-to-image model, built for high-resolution output, strong prompt adherence, and fine-grained typography — legible text inside the image, including Chinese characters.ChromaImage
Chroma is an open-weight, Apache-2.0 text-to-image model by Lodestone — a true base model with no aesthetic tuning, meant as a raw, neutral foundation for fine-tuning.QwenImage
Qwen-Image is an open-weight text-to-image model from Alibaba, built for sharp prompt adherence and unusually legible in-image text — including strong results with Chinese characters.Stable DiffusionImage
Stable Diffusion is the open model that kicked off the AI-art community — and its 1.5 release is still the fastest, lightest, and most widely supported version to generate with.HiDreamImage
HiDream I1 is an open-weight text-to-image model from HiDream.ai built on a 17B sparse mixture-of-experts transformer, known for strong prompt adherence and legible in-image text.Krea 2NewImage
Krea 2 is Krea AI's in-house text-to-image model, built for sharp photorealism, strong aesthetics, and dependable prompt adherence.AnimaImage
Anima is an open-weight anime and illustration model from CircleStone Labs, built for clean linework, expressive characters, and vivid color straight from a prompt.Z-ImageImage
Z-Image is a family of open-weight text-to-image models from Alibaba's Tongyi Lab, built on a compact ~6B architecture that punches well above its size on prompt adherence, clean composition, and legible in-image text — including Chinese and English.Video models
WanVideo
Wan is a family of open-weight video models from Alibaba that turn a text prompt or a still image into short, cinematic clips with fluid, realistic motion.LTX VideoVideo
LTX Video (LTXV) is an open-weight video model from Lightricks built around a diffusion transformer tuned for speed — it generates coherent, cinematic clips from a text prompt or a starting image at close to real-time rates.KlingVideo
Kling is a family of high-end text-to-video and image-to-video models from Kuaishou, one of China's largest short-form video platforms.SeedanceVideo
Seedance is ByteDance's hosted multimodal video model that turns a text prompt or a still image into a short clip — and generates the audio with it, producing synchronized dialogue, lip-sync, and ambient sound in a single pass.Grok ImagineVideo
Grok Imagine is xAI's hosted model for turning a text prompt or a reference image into short animated video clips — and it generates still images too.HappyHorseVideo
HappyHorse is a hosted video model that turns a text prompt — or a single still image — into a short clip, and it generates the matching sound in the same pass, so effects like a splashing wave or engine noise line up with the on-screen action.Veo 3Video
Veo 3 is Google DeepMind's state-of-the-art AI video model, and its headline feature is native audio — it generates dialogue, sound effects, and even music together with the picture, from a text prompt or a reference image.Sora 2Video
Sora 2 is OpenAI's text-to-video generation model, built for physically plausible motion, coherent multi-shot scenes, and synchronized audio generated alongside the video.