Which Qwen-Image.
Four things go by the name. Three of them share one text encoder and one VAE.
- Qwen-Image, from 4 August 2025, is the original: a 20B model under Apache 2.0.
- Qwen-Image 2512, from 31 December 2025, is the “December update” with better human realism and text. Same size, same encoder, same VAE. Users say it lost some of the flat anime look of the August model.
- Lightning LoRAs from lightx2v cut either one to 8 or 4 steps at CFG 1. There are separate LoRAs for the original and for 2512.
- Qwen-Image 2.1, from 20 September 2026, is a new model: 7.1B, with a Qwen3-VL 8B encoder, its own VAE with transparency, and text to image plus editing from up to 10 reference images in one file. Native size is 2048 × 2048.
The editing sibling, Qwen-Image-Edit (2509 and 2511), is a separate model. It gets a short section under the files.
Files you need.
One model, one text encoder, one VAE. Qwen-Image and 2512 take Qwen2.5-VL 7B. Qwen-Image 2.1 takes Qwen3-VL 8B. They don’t swap.
Qwen-Image and 2512
-
Model20.4 GB Download
qwen_image_fp8_e4m3fn.safetensorsComfyUI/models/diffusion_models/ 2512: qwen_image_2512_fp8_e4m3fn.safetensors, same size. Full precision: 40.9 GB -
Text encoder9.4 GB Download
qwen_2.5_vl_7b_fp8_scaled.safetensorsComfyUI/models/text_encoders/ full precision: qwen_2.5_vl_7b.safetensors, 16.6 GB -
VAE0.3 GB Download
qwen_image_vae.safetensorsComfyUI/models/vae/
Lightning LoRAs
Use the LoRA made for your model, and match the step count to the file: lightx2v says 4-step LoRAs technically run at 8 steps, but the 8-step LoRA at 8 and the 4-step LoRA at 4 is the good choice.
-
Qwen-Image, 8 steps0.9 GB Download
Qwen-Image-Lightning-8steps-V2.0-bf16.safetensorsComfyUI/models/loras/ -
Qwen-Image, 4 steps0.9 GB Download
Qwen-Image-Lightning-4steps-V2.0-bf16.safetensorsComfyUI/models/loras/ -
2512, 8 steps0.9 GB Download
Qwen-Image-2512-Lightning-8steps-V1.0-bf16.safetensorsComfyUI/models/loras/ -
2512, 4 steps0.9 GB Download
Qwen-Image-2512-Lightning-4steps-V1.0-bf16.safetensorsComfyUI/models/loras/
Each also comes as a 1.7 GB fp32 file. lightx2v also has merged 2512 checkpoints with the LoRA built in, 20.4 GB each.
Qwen-Image 2.1
-
Model7.3 GB Download
qwen_image_2.1_int8_convrot.safetensorsComfyUI/models/diffusion_models/ full precision: qwen_image_2.1_bf16.safetensors, 14.2 GB -
Text encoder9.4 GB Download
qwen3vl_8b_int8_convrot.safetensorsComfyUI/models/text_encoders/ full precision: qwen3vl_8b_bf16.safetensors, 17.5 GB -
VAE0.7 GB Download
qwen_image_2.1_vae_bf16.safetensorsComfyUI/models/vae/
The same repo has two qwen3.5_9b_qwen_image_2.1_pe files. They are prompt enhancers, language models for ComfyUI’s TextGenerate node, not text encoders. Loading one as the encoder gives garbage or a shape error. The int8 files run fast only on PyTorch built for CUDA 13.0 (cu130); on cu128 they were slower than a GGUF in Kijai’s test.
models/
├── diffusion_models/
│ ├── qwen_image_fp8_e4m3fn.safetensors
│ └── qwen_image_2.1_int8_convrot.safetensors
├── loras/
│ └── Qwen-Image-Lightning-8steps-V2.0-bf16.safetensors
├── text_encoders/
│ ├── qwen_2.5_vl_7b_fp8_scaled.safetensors
│ └── qwen3vl_8b_int8_convrot.safetensors
└── vae/
├── qwen_image_vae.safetensors
└── qwen_image_2.1_vae_bf16.safetensors
Smaller files: GGUF
Qwen-Image and 2512 GGUFs load through the ComfyUI-GGUF nodes and go in diffusion_models. For the encoder, unsloth’s Qwen2.5-VL 7B GGUF is 4.7 GB at Q4_K_M and 8.1 GB at Q8. unsloth also has Qwen-Image 2.1 GGUFs (4.2 GB at Q4_K_M), but ComfyUI-GGUF’s main branch doesn’t load them yet.
| Model | Q8_0 | Q5_K_M | Q4_K_M | Q4_0 |
|---|---|---|---|---|
| Qwen-Image | 21.8 GB | 14.9 GB | 13.1 GB | 11.9 GB |
| 2512 | 21.8 GB | 15.0 GB | 13.2 GB | 11.9 GB |
city96 goes lower too: Q3_K_M is 9.7 GB and Q2_K 7.1 GB.
Qwen-Image-Edit
Edit is its own model for changing an existing picture. It uses the same Qwen2.5-VL 7B encoder and Qwen-Image VAE as above, plus its own file from Comfy-Org’s Edit repo: qwen_image_edit_2511_fp8mixed.safetensors is the newest, 20.5 GB. In ComfyUI the picture goes in through the TextEncodeQwenImageEditPlus node, which takes the prompt and up to three images. If you use a GGUF encoder for Edit, it needs its matching mmproj file next to it, or the image input breaks. The ComfyUI Edit tutorial has the full walkthrough. Qwen-Image 2.1 can also edit, from reference images, with the files above.
What fits your computer.
At about 1 megapixel. The text encoder runs first and ComfyUI moves it out of the way before sampling, so the model file sets the limit. A model that doesn’t fit still runs with part of it in system RAM, only slower.
- 6 GBNo
Even Qwen-Image Q2_K is 7.1 GB. Z-Image Turbo or Flux.2 Klein 4B are the better fit.
- 8 GBOffloads
Only with most of the model in system RAM. Qwen-Image 2.1 int8 (7.3 GB) comes closest. Plan on 32 GB of RAM or more.
- 12 GBOffloads
Qwen-Image as GGUF Q4_0 (11.9 GB) or Q3_K_M (9.7 GB), partly in RAM. Qwen-Image 2.1 int8 fits better.
- 16 GBFits
Qwen-Image 2.1 int8 with the int8 encoder: an RTX 4060 Ti 16 GB peaked around 15 GB with no offloading. Qwen-Image as Q4_K_M or Q5_K_M GGUF, or fp8 with offloading and the Lightning LoRA.
- 24 GBFits
Qwen-Image or 2512 in fp8. Comfy’s own timings are from a 24 GB RTX 4090D at 86% of its memory.
- Mac 24 GBSwaps
Qwen-Image 2.1 swaps hard: one edit with three references at 1280 × 736 took 35 minutes on an M4 Pro with 24 GB.
- Mac 32 to 48 GBTight
Qwen-Image as GGUF (13.1 to 21.8 GB) with a GGUF encoder, or Qwen-Image 2.1 in bf16, about 32 GB for the set. We found no measured Mac times to quote.
- Mac 64 GB+Fits
Qwen-Image 2.1 in bf16 with room to spare, or Qwen-Image as Q8_0 GGUF.
Set it up.
-
Update ComfyUI
Qwen-Image needs a ComfyUI from August 2025 or later; older ones say
qwen_imageis not in the list of CLIP types. Qwen-Image 2.1 needs ComfyUI 0.37.0 or the nightly build, for itsTextEncodeQwenImage21node. ComfyUI Desktop and the portable build can lag behind. In a manual install:Terminal, in the ComfyUI foldergit pull pip install -r requirements.txt
-
Download the files
Model, text encoder and VAE from the lists above, plus a Lightning LoRA if you want 8 or 4 steps.
-
Put them in their folders
Model in
diffusion_models, encoder intext_encoders, VAE invae, LoRA inloras. Restart ComfyUI so the loaders list them. -
Open a Qwen template
In the template browser, pick Qwen-Image: Text to Image, Qwen Image 2512 or the Qwen-Image 2.1 text to image template. The first two ship with a Lightning LoRA already wired in.
-
Check the loaders
Load Diffusion Model gets the model. Load CLIP gets the encoder with type
qwen_image. Load VAE gets the matching VAE. For Qwen-Image and 2512 there is aModelSamplingAuraFlownode at shift 3.1; for 2.1 the prompt goes throughTextEncodeQwenImage21instead of a plain text encode. -
Write a prompt and run
Plain, detailed sentences work. The first run loads everything and takes longest.
Settings that work.
Qwen-Image and 2512
- Steps
- 50
- CFG
- 4
- Sampler
- euler
- Scheduler
- simple
- Size
- 1328 × 1328
- Shift
- AuraFlow 3.1
- Negative
- yes
- Latent
- EmptySD3LatentImage
That is the model card: 50 steps at a true CFG of 4. Comfy’s first Qwen-Image template ran 20 steps at CFG 2.5, which is much faster and still good; its note says to try 50 for the original settings. The 2512 template uses 50 steps at CFG 4.
With a Lightning LoRA
- Steps
- 8 or 4
- CFG
- 1
- Sampler
- euler
- Scheduler
- simple
Steps match the LoRA’s name. Comfy’s current Qwen-Image template runs the 8-step LoRA at 8 steps, CFG 1, shift 3.1. Don’t combine Lightning with the community qwen_image_distill_full model: that one is already distilled, and runs at 15 steps and CFG 1 on its own.
| Aspect | Size |
|---|---|
| 1:1 | 1328 × 1328 |
| 16:9 | 1664 × 928 |
| 9:16 | 928 × 1664 |
| 4:3 | 1472 × 1140 |
| 3:4 | 1140 × 1472 |
| 3:2 | 1584 × 1056 |
| 2:3 | 1056 × 1584 |
The model card’s sizes. Comfy’s 2512 note uses 1472 × 1104 for 4:3, a multiple of 16. Users report anything from 768 to 2048 on a side works well.
Qwen-Image 2.1
- Steps
- 25 to 40
- CFG
- 1
- Sampler
- euler
- Scheduler
- simple
- Size
- 1024 to 2048
- Multiple of
- 32
- Text encode
- TextEncodeQwenImage21
- Latent
- EmptyLatentImage
The model card uses 40 steps at 2048 × 2048. Comfy’s template uses 25 steps at 1024 × 1024, which is quicker; users who saw faint banding fixed it with 40 steps at 2K. Keep CFG at 1 unless you add a negative prompt, then raise it. For a transparent PNG, wrap the prompt the way the model card does: start with “This is an RGBA image with transparency.” and end with “The image has alpha channel and the background is transparent.”
How fast.
| GPU | Model | Steps | Time |
|---|---|---|---|
| RTX 4090D 24 GB | Qwen-Image fp8 | 20 | 71 s[1] |
| RTX 4090D 24 GB | fp8 + 8-step Lightning | 8 | 34 s[1] |
| RTX 5070 Ti 16 GB | 2512 GGUF Q5_K_M, 1024 px | 40 | 148 s[2] |
| RTX 4060 Ti 16 GB | 2.1 int8, text to image | not given | 20 s[3] |
| RTX 4060 Ti 16 GB | 2.1 int8, edit | not given | 60 s[3] |
Times after the first run, which loads the models (the 4090D took 94 s and 55 s on its first runs).
On an RTX 20-series card it is a different story: those GPUs have no bf16, fp16 overflows to black, so Qwen-Image runs in fp32 and one image took more than 10 minutes on a 2080 Ti.
When it goes wrong.
CLIPLoader ... Value not in list: type: 'qwen_image' not in [...]- ComfyUI is too old for Qwen-Image. Update it.
- The image turns black partway through sampling
- Start ComfyUI without
--use-sage-attention, and don’t set the weight dtype tofp8_e4m3fn_fast. The same fix applies to black Lightning images. Error while deserializing header: header too large- A broken download, usually the VAE. Download it again.
lora key not loaded, or blurry low-contrast Lightning images- Your ComfyUI is older than the LoRA support. Update it.
- An error with a Qwen2.5-VL file from elsewhere
- Wrong text encoder. Use Comfy-Org’s
qwen_2.5_vl_7b_fp8_scaledor the full-precision one from the same repo. TextEncodeQwenImage21orQwenImage21Cachemissing- Qwen-Image 2.1 needs ComfyUI 0.37.0 or the nightly build.
Given normalized_shape=[4096] ... got input of size [1, 338, 5120]- Qwen-Image 2.1 got the wrong encoder, often a
qwen3.5_9bprompt-enhancer file. Useqwen3vl_8b_int8_convrotorqwen3vl_8b_bf16. - Qwen-Image 2.1 looks yellow, or a fine diamond grid shows on skin
- For the tint, raise CFG and add a negative prompt. The grid comes from the 2.1 VAE; madebyollin’s texture-fix VAE is a drop-in replacement.
With the 4-step Lightning LoRA every seed gives nearly the same picture. Is there a way to get more variety?
The licence says non-commercial, the team says outputs are fine. Which is it for Qwen-Image 2.1?
For the first, users VAE-encode a blank image and sample it at a denoise of about 0.95 instead of starting from an empty latent. The second has no official answer yet.
Questions.
Which text encoder does Qwen-Image use?
Qwen-Image and Qwen-Image 2512 use Qwen2.5-VL 7B, loaded with Load CLIP set to type qwen_image. Qwen-Image 2.1 uses Qwen3-VL 8B. The qwen3.5_9b files in the 2.1 repo are prompt enhancers, not encoders.
How much VRAM does Qwen-Image need?
24 GB runs the 20.4 GB fp8 model without offloading. 16 GB works with a Q4_K_M or Q5_K_M GGUF, or fp8 with part of it in system RAM. Qwen-Image 2.1 is smaller: its int8 files peaked around 15 GB on a 16 GB card.
Can I use Qwen-Image commercially?
Qwen-Image and Qwen-Image 2512 are Apache 2.0, so yes. Qwen-Image 2.1 is under the Qwen Research License, for non-commercial purposes only, with a separate commercial licence from Qwen.
What is the difference between Qwen-Image and 2512?
2512 is the December 2025 update of the same model, with better human realism and text. It uses the same encoder and VAE and has the same file size. Some users prefer the original for flat anime styles.
Which Lightning LoRA should I use?
The one made for your model: Qwen-Image-Lightning for the original, Qwen-Image-2512-Lightning for 2512. Run the 8-step LoRA at 8 steps and the 4-step one at 4, both at CFG 1.
Do Qwen-Image 2.1 GGUFs work in ComfyUI?
Not with the main branch of ComfyUI-GGUF yet; support is in open pull requests. Use the int8 files (7.3 GB model, 9.4 GB encoder) for now. Qwen-Image and 2512 GGUFs work.
How do I edit images with Qwen?
Either with Qwen-Image-Edit 2511, a separate model that shares Qwen-Image’s encoder and VAE, through ComfyUI’s Qwen-Image-Edit templates. Or with Qwen-Image 2.1, which edits from reference images with the same file it uses for text to image.
Sources: Qwen-Image model card, Qwen-Image 2512 model card, Qwen-Image 2.1 model card, ComfyUI Qwen-Image tutorial, ComfyUI Qwen-Image 2.1 tutorial, Comfy Qwen-Image template [1], 2512 speed thread [2], Qwen-Image 2.1 on a 4060 Ti [3], Qwen-Image-Lightning, ComfyUI issue #10852, ComfyUI issue #16470, 2.1 on an M4 Pro.