Kirazuri (Ernie)
Kirazuri (Ernie) is a full fine-tune of the ERNIE-Image model.
ERNIE-Image is an open text-to-image generation model developed by the ERNIE-Image team at Baidu.
Version 0.1 (Latest) is trained on ~50,000 images at multiple resolutions in three total stages up to a maximum resolution of 1024^2 pixels.
This fine-tune focuses on several goals:
- Learn new concepts/styles/characters associated with tag-based prompting.
- Enhance the model aesthetic guided by manually applied quality, aesthetic, and style tagging
- Preserve the text-rendering capabilities in multiple languages (English, Chinese, Japanese)
For more details, see the Kirazuri (Ernie) Training Diary
Training Details Summary
Trainer: diffusion-pipe
Training device: NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition
Total training time: ~500 hrs (~20 days)
Total samples seen (un-batched steps): ~1,328,000
Training resolutions: 512^2, 768^2, 1024^2
Additional Features
- Tag Dropout: 10% with protected first 8 tags
- Tag Shuffle: Applied to last unprotected tags
- Natural Language: 6 total Short and Long Caption variants
Quantizations
Int8 convrot quants of the model and text encoder are available and strongly recommended for use in ComfyUI.
Generation speed *(on NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition):
Before:
26gb vram, 1.35it/s
After:
18gb vram, 1.89it/s
Installing and running
Workflow:

The model is natively supported in ComfyUI. The above image contains a workflow; you can open it in ComfyUI or drag-and-drop to get the workflow.
Generation Settings
Recommendations:
1024^2 Resolutions
50 Steps
CFG 4
er_sde or euler Sampler
*Some previews are generated at 1280^2 resolutions - despite no training yet at this resolution, the model can still perform well for small details here.
*The model can converge in as few as 20 steps, but anatomy and prompt adherence will suffer.
Prompting
This model is trained on combinations of booru-style tags and natural language captions.
It does not perform well yet with tag-only prompts, and they are best used in combination with long natural language descriptions.
This follows from the base models natural language only training, and intended use with a prompt enhancer that expands descriptions to long-format text.
Tag order
Tags combined with natural language can be placed at the start or end of the prompt.
[quality/meta/safety tags] [character] [series] [artist] [1girl/1boy/1other etc] [general tags]
[quality/meta/safety tags] [character] [series] [artist] tag groups are also not shuffled, so their order may have some influence on generations.
Quality and Aesthetic tags
Human score based: masterpiece, best quality, very aesthetic, aesthetic
Meta tags
absurdres, official art, etc
Styles
painterly, chiaroscuro, ligne claire, flat color, no lineart, blending, etc
traditional media, oil painting \(medium\), watercolor \(medium\), etc
Recognitions
Thanks to Baidu labs for the open research and release of the ERNIE-Image model.
Thanks to tdrussell of CircleStone Labs for the diffusion-pipe trainer.
Thanks to narugo1992 and the deepghs team for open-sourcing various training sets, image processing tools, and models.
Thanks to silveroxides for the convert_to_quant quantization tool.
License
Apache license 2.0
Note that ERNIE-Image is governed by the same license; this fine-tune does not grant any rights or impose any restrictions to the base weights.




