Sign In

Krea 2 Turbo optimized for AMD ROCm

28

Download

1 variant available

int8 SafeTensor

krea2_turbo_style_ref_fused_int8_tensorwise_v1.safetensors

8-bit integer, smaller file • 13.16 GB

Verified:

Type
Checkpoint Trained
Stats

32

Reviews
Published

Sep 30, 2026

Base Model

Krea 2

Hash
AutoV2
32189D0FF9
default creator card background decoration
Followers - 125

125

Likes - 116

116

License:

krea_2_turbo_00009_.png

Krea 2 Turbo INT8 TensorWise — AMD ROCm Optimized for ComfyUI

⚠️ Running the models alone is not enough to unlock full AMD optimization. Follow all implementation instructions below to build a fully AMD ROCm-optimized ComfyUI stack.

⚡ Krea 2 Turbo optimized for AMD ROCm, ComfyUI, FlashAttention, Triton/AITER, and low-memory inference.

Release video


This release is the result of a full Krea 2 optimization project focused on making Krea 2 Turbo and Style Reference substantially more practical on AMD GPUs.

It combines:

  • Krea 2 Turbo INT8 TensorWise diffusion

  • Krea 2 Turbo Style Reference with the LoRA pre-fused before INT8 quantization

  • Krea Qwen3-VL-4B INT8 TensorWise text encoder

  • ROCm FlashAttention using AMD Triton/AITER

  • Qwen VAE Triton W8A8 acceleration

  • optimized Text-to-Image and Reference-to-Image ComfyUI workflows

The diffusion checkpoints use native ComfyUI int8_tensorwise quantization and do not depend on ConvRot.


🚀 Performance

On the same Ubuntu AMD ROCm host, using a 1368×768 / 8-step reference-to-image workflow:

ConfigurationFirst workflowWarm workflowWarm samplingStock-style ComfyUI INT8 + runtime Style Reference LoRA + PyTorch cross-attention242.4s81.7s71.0sFully optimized project stack85.0s75.0s65.0s

Measured gains

  • First workflow: 242.4s → 85.0s

    • 64.9% lower end-to-end time

  • Warm workflow: 81.7s → 75.0s

    • 8.2% faster

  • Warm sampling: 71.0s → 65.0s

    • 8.4% faster

  • Iteration time: 8.87s/it → 8.16s/it

    • 8.0% lower

Cold-start results include initialization and compilation overhead. Warm sampling and sec/it are better indicators of steady-state inference performance.

FlashAttention matters

On the tested system, switching from PyTorch cross-attention to ROCm FlashAttention with AMD Triton/AITER reduced:

  • warm generation: 82.5s → 75.0s

  • warm sampling: 71.0s → 65.0s

  • warm sec/it: 8.915 → 8.160

ComfyUI's separate --enable-triton-backend flag produced no meaningful additional gain in this particular A/B test, so it is not required for this release.


📦 Which files do I need?

The full release contains three optimized model components.

1. Krea 2 Turbo INT8 TensorWise

krea2_turbo_int8_tensorwise_v1.safetensors

Use this for normal Krea 2 Turbo text-to-image generation.

  • all 224 main transformer GEMM weights converted to INT8 TensorWise

  • no ConvRot

  • approximately 13.16 GiB

  • approximately 46% smaller than the original BF16 transformer

Place it in:

ComfyUI/models/diffusion_models/

2. Krea 2 Turbo Style Reference — Fused INT8 TensorWise

krea2_turbo_style_ref_fused_int8_tensorwise_v1.safetensors

Use this for Krea 2 Style Reference / reference-to-image workflows.

The Ostris Style Reference LoRA has already been fused into the BF16 Krea 2 Turbo model before quantization.

That means:

✅ No runtime Style Reference LoRA loading
✅ 224 fused transformer weights quantized to INT8 TensorWise
✅ 32 fused txtfusion weights retained in BF16
✅ Reference conditioning remains available

Important

Do not load the Style Reference LoRA again.

The LoRA is already fused into this checkpoint.

You still need the Krea 2 Ostris Edit/reference-conditioning path:

https://github.com/ostris/ComfyUI-Krea2-Ostris-Edit

Original Style Reference project:

https://huggingface.co/ostris/krea2_turbo_style_reference

Place the fused model in:

ComfyUI/models/diffusion_models/

3. Krea Qwen3-VL-4B INT8 TensorWise Text Encoder

krea2_qwen3vl_4b_int8_tensorwise_v1.safetensors

This converts 356 attention, MLP, and vision projection weights to native ComfyUI INT8 TensorWise.

Final size is approximately 4.50 GiB, around 45% smaller than the verified BF16 source.

Place it in:

ComfyUI/models/text_encoders/

Load it using ComfyUI CLIPLoader with:

type: krea2
device: default

⚡ Recommended AMD ROCm setup

For the best performance measured during this project, use:

  • AMD ROCm-capable GPU

  • current compatible ROCm + PyTorch environment

  • ROCm/flash-attention

  • AMD Triton/AITER FlashAttention backend

  • Qwen VAE Triton W8A8

  • the optimized models from this release

Set:

export FLASH_ATTENTION_TRITON_AMD_ENABLE=TRUE

Launch ComfyUI with:

python main.py \
  --use-flash-attention \
  --disable-xformers

Do not blindly copy ROCm, PyTorch, Triton, GFX override, or environment-variable settings from another AMD machine.

GPU architecture and ROCm support should be detected first.


🧠 Copy/Paste AMD ROCm Installation Prompt

I created a complete installation prompt that can be pasted into ChatGPT, Claude, Gemini, or another capable LLM with web access.

It guides you interactively through:

  • GPU and GFX architecture discovery

  • AMD driver validation

  • ROCm/HIP setup

  • Python 3.12 environment creation

  • compatible PyTorch + TorchVision selection

  • ROCm FlashAttention installation

  • AITER and Triton validation

  • ComfyUI installation

  • environment-variable tuning

  • final GPU and generation validation

ROCm FlashAttention + AMD Triton/AITER installation prompt:

https://huggingface.co/PuppetVision/krea-2-amd-rocm-optimized-comfy-triton#copy-and-paste-implementation-instructions-for-chatgpt-or-other-flagship-llm


⚡ Qwen VAE Triton W8A8

Krea 2's Qwen Image VAE can be a major cold-start and resolution-change bottleneck.

My Qwen VAE Triton W8A8 custom node accelerates validated VAE decoder layers while keeping quality-sensitive layers in native precision.

Docker isolation benchmarks showed approximately 20–21% lower VAE execution cost in the tested first-run and resolution-change cases.

GitHub:

https://github.com/AllenCraigBarnard/ComfyUI-Qwen-VAE-Triton

ComfyUI Manager search:

Qwen VAE Triton W8A8

Recommended preset:

Aggressive

Connect it as:

VAELoader
   ↓
Patch Qwen VAE Triton W8A8
   ↓
VAEDecode

Use the normal:

qwen_image_vae.safetensors

The VAE itself is not included in this model release.


🧩 Optimized ComfyUI Workflows

Validated Text-to-Image and Style-Reference workflows are available here:

https://huggingface.co/PuppetVision/krea-2-amd-rocm-optimized-comfy-triton/workflows

Full Hugging Face release:

https://huggingface.co/PuppetVision/krea-2-amd-rocm-optimized-comfy-triton

The Hugging Face repository also contains:

  • benchmarks

  • validation evidence

  • model provenance

  • tensor policies

  • checksums

  • licenses

  • optimized workflows

  • detailed AMD ROCm installation guidance


🛠️ Basic Installation

Place the optimized models here:

ComfyUI/
└── models/
    ├── diffusion_models/
    │   ├── krea2_turbo_int8_tensorwise_v1.safetensors
    │   └── krea2_turbo_style_ref_fused_int8_tensorwise_v1.safetensors
    │
    └── text_encoders/
        └── krea2_qwen3vl_4b_int8_tensorwise_v1.safetensors

For the optimized VAE path, also install:

https://github.com/AllenCraigBarnard/ComfyUI-Qwen-VAE-Triton

For Style Reference, install:

https://github.com/ostris/ComfyUI-Krea2-Ostris-Edit


🔬 Technical Notes

Diffusion quantization

The primary Krea 2 Turbo model contains 28 main transformer blocks.

Eight large GEMM families per block are quantized:

attn.wq
attn.wk
attn.wv
attn.gate
attn.wo
mlp.gate
mlp.up
mlp.down

Total:

28 × 8 = 224 INT8 TensorWise weights

Each quantized layer uses native ComfyUI TensorWise metadata:

.weight        INT8
.weight_scale  FP32
.comfy_quant   {"format":"int8_tensorwise"}

Style Reference fusion

The Style Reference LoRA was fused at multiplier 1.0 while the base model was still BF16:

W_fused = W_base + (B @ A)

The fused transformer weights were then quantized.

This avoids:

INT8 → dequantize → merge LoRA → requantize

⚠️ Compatibility

This project was developed and benchmarked primarily for AMD ROCm + ComfyUI.

Performance varies according to:

  • GPU architecture

  • ROCm version

  • PyTorch version

  • FlashAttention/AITER/Triton versions

  • ComfyUI revision

  • resolution

  • workflow

  • memory configuration

  • kernel compilation state

Benchmark numbers are measurements from the tested environments, not universal guarantees.


📜 License & Attribution

This is an independent community derivative release.

It is not an official Krea, Qwen/Alibaba, Ostris, ComfyUI, AMD, or KJnodes product, and it is not endorsed by those parties.

Krea 2 Turbo / Style Reference

Krea 2 Turbo and its derivatives are governed by the Krea 2 Community License Agreement.

Please review the license before redistribution or commercial deployment:

https://www.krea.ai/krea-2-licensing

Official Krea 2 Turbo:

https://huggingface.co/krea/Krea-2-Turbo

Qwen3-VL

The upstream Qwen3-VL-4B-Instruct text encoder is published under Apache License 2.0:

https://huggingface.co/Qwen/Qwen3-VL-4B-Instruct

The Hugging Face release contains the complete license copies and modification/attribution notices.


🚀 More Projects / Contact

If this project saved you VRAM, inference time, or days of debugging an AMD Krea 2 setup, check out my other optimization work.

LinkedIn — Allen Barnard
https://www.linkedin.com/in/allen-b-3a35505a/

YouTube — PuppetVisionAI
https://www.youtube.com/@PuppetVisionAI

Website
https://puppetvision.nl

GitHub / More Projects
https://github.com/AllenCraigBarnard/ComfyUI-Qwen-VAE-Triton#work-with-me--more-projects

I’m interested in AI Systems Engineering opportunities involving model optimization, inference systems, GPU acceleration, quantization, ROCm/CUDA, Triton, and generative-AI infrastructure.


Full model release:
https://huggingface.co/PuppetVision/krea-2-amd-rocm-optimized-comfy-triton

Optimized workflows:
https://huggingface.co/PuppetVision/krea-2-amd-rocm-optimized-comfy-triton/workflows

ROCm FlashAttention / AMD Triton installation prompt:
https://huggingface.co/PuppetVision/krea-2-amd-rocm-optimized-comfy-triton/docs/ROCM_FLASH_ATTENTION_INSTALLATION_PROMPT.md

Qwen VAE Triton W8A8:
https://github.com/AllenCraigBarnard/ComfyUI-Qwen-VAE-Triton

Release video:
https://youtu.be/fD-J2CYetIg