Sign In

ComfyUI Qwen image 2.1 Enhancer

Updated: Oct 7, 2026

toolqwen image 2.1

Download

1 variant available

Archive Other

ComfyUI-qwen_img_2_1_enhancer-main.zip

3.01 MB

Verified:

Type
Other
Stats

148

Reviews
Published

Oct 4, 2026

Base Model

Qwen 2.1

Hash
AutoV2
70CD46FBE3
default creator card background decoration
Followers - 646

646

Likes - 2863

2.9K

Downloads - 55561

55.6K

Qwen is licensed under the Qwen RESEARCH LICENSE AGREEMENT, Copyright (c) 2026 Hangzhou Tongyi Laboratory Technology Co., Ltd. All Rights Reserved.

Screenshot from 2026-10-04 00-08-33.png

v1.1.0

Follow up update of this post here

I added mask support to the reference strength node. You can now increase or reduce attention to a selected part of a reference instead of adjusting the whole photo. also the photo in the shown workflow is missing the rest of the outfit due to the masking mode I chose and how tight I masked it as this was deliberate.

Connect the node between your model loader and sampler, connect your reference's mask, and select its image index: 1 for the first connected reference, 2 for the second, etc. Keep your images and VAE connected to the encoder as usual.

There are two modes:

- focus_only: applies strength to the selected tokens and leaves the rest at native weighting.

- zero_unmasked_tokens: also blocks direct attention to the unselected reference tokens throughout the diffusion transformer.

mask_threshold controls how much of a token the mask must cover: 1 requires full coverage, lower values include more edge tokens, and 0 selects everything. strength at 1 is native, above 1 increases priority, and below 1 reduces it.

Match mask_resize_method to your image resizing: lanczos when the native encoder resizes the original, or nearest-exact for an external nearest-exact resize. Chain nodes for separate references.

Still a work in progress. This controls reference attention; it isn't an output-area lock. The encoders still process the full image, so blocking a reference token doesn't erase information already carried into other tokens or conditioning.

Download and detailed usage on GitHub

Sample workflow here


I put together an enhancer pack for Qwen Image 2.1 with two nodes for controlling what the model pays attention to during editing.

Reference Strength lets you select a reference image and increase or decrease its attention priority. If you're working with multiple references, you can adjust them separately instead of giving every image the same treatment, and you also can use it for single image to prevent the loss of likeliness at times.


ref index starts from 1; meaning image_1 and same for the rest of images, where image_2 is ref index 2 in the node. ( Soon adding mask support)

Phrase Weights brings phrase-level attention control to the Qwen Image 2.1 edit encoder. So you can write:

Add (warm sunset lighting:1.4) with (soft shadows:1.2).

and give those specific parts of the prompt more attention while leaving the rest at its normal weighting.

It works with both positive and negative prompts, with separate weights for each. There's also an inspection output showing exactly which token rows and pieces were matched.

For both controls:

`1.0` = untouched  
`>1.0` = more attention priority  
`<1.0` = less attention priority  
`0.0` = suppression  

The adjustments happen inside Qwen's attention during sampling. The prompt and reference images still go through the native encoding path, without multiplying the finished text embeddings or pasting reference pixels into the output.

You can use either node on its own (I prefer this), or combine them. Reference strength applies to the whole selected image, and higher weights can make a reference or phrase dominate, so there's still some balancing to do.

More examples can be found here : https://www.reddit.com/r/StableDiffusion/comments/1wx6t0v/comfyui_qwen_image_21_enhancer_two_nodes/

Installation, usage, and the technical details are in the repo:
ComfyUI-qwen_img_2_1_enhancer

Ref_strength_workflow

Phrase_Weights_workflow