Image model

How to run Qwen-Image locally.

Qwen-Image runs in ComfyUI with three files: the 20.4 GB fp8 model, the Qwen2.5-VL 7B text encoder and the Qwen-Image VAE. It runs without offloading on a 24 GB card, and as a GGUF with part of it in system RAM on 12 to 16 GB. Qwen-Image 2.1 is a smaller, newer model with its own encoder, and a non-commercial licence.

Updated 29 Sep 202611 min read

Maker
Qwen team, Alibaba
Released
Aug 20252512 in Dec 2025, 2.1 in Sep 2026
Licence
Apache 2.02.1: non-commercial research licence
Memory
16 GB and up24 GB for fp8 without offloading

Which Qwen-Image.

Four things go by the name. Three of them share one text encoder and one VAE.

  • Qwen-Image, from 4 August 2025, is the original: a 20B model under Apache 2.0.
  • Qwen-Image 2512, from 31 December 2025, is the “December update” with better human realism and text. Same size, same encoder, same VAE. Users say it lost some of the flat anime look of the August model.
  • Lightning LoRAs from lightx2v cut either one to 8 or 4 steps at CFG 1. There are separate LoRAs for the original and for 2512.
  • Qwen-Image 2.1, from 20 September 2026, is a new model: 7.1B, with a Qwen3-VL 8B encoder, its own VAE with transparency, and text to image plus editing from up to 10 reference images in one file. Native size is 2048 × 2048.

The editing sibling, Qwen-Image-Edit (2509 and 2511), is a separate model. It gets a short section under the files.

Files you need.

One model, one text encoder, one VAE. Qwen-Image and 2512 take Qwen2.5-VL 7B. Qwen-Image 2.1 takes Qwen3-VL 8B. They don’t swap.

Qwen-Image and 2512

  • Model qwen_image_fp8_e4m3fn.safetensors ComfyUI/models/diffusion_models/ 2512: qwen_image_2512_fp8_e4m3fn.safetensors, same size. Full precision: 40.9 GB
    20.4 GB Download
  • Text encoder qwen_2.5_vl_7b_fp8_scaled.safetensors ComfyUI/models/text_encoders/ full precision: qwen_2.5_vl_7b.safetensors, 16.6 GB
    9.4 GB Download
  • VAE qwen_image_vae.safetensors ComfyUI/models/vae/
    0.3 GB Download

Lightning LoRAs

Use the LoRA made for your model, and match the step count to the file: lightx2v says 4-step LoRAs technically run at 8 steps, but the 8-step LoRA at 8 and the 4-step LoRA at 4 is the good choice.

  • Qwen-Image, 8 steps Qwen-Image-Lightning-8steps-V2.0-bf16.safetensors ComfyUI/models/loras/
    0.9 GB Download
  • Qwen-Image, 4 steps Qwen-Image-Lightning-4steps-V2.0-bf16.safetensors ComfyUI/models/loras/
    0.9 GB Download
  • 2512, 8 steps Qwen-Image-2512-Lightning-8steps-V1.0-bf16.safetensors ComfyUI/models/loras/
    0.9 GB Download
  • 2512, 4 steps Qwen-Image-2512-Lightning-4steps-V1.0-bf16.safetensors ComfyUI/models/loras/
    0.9 GB Download

Each also comes as a 1.7 GB fp32 file. lightx2v also has merged 2512 checkpoints with the LoRA built in, 20.4 GB each.

Qwen-Image 2.1

  • Model qwen_image_2.1_int8_convrot.safetensors ComfyUI/models/diffusion_models/ full precision: qwen_image_2.1_bf16.safetensors, 14.2 GB
    7.3 GB Download
  • Text encoder qwen3vl_8b_int8_convrot.safetensors ComfyUI/models/text_encoders/ full precision: qwen3vl_8b_bf16.safetensors, 17.5 GB
    9.4 GB Download
  • VAE qwen_image_2.1_vae_bf16.safetensors ComfyUI/models/vae/
    0.7 GB Download

The same repo has two qwen3.5_9b_qwen_image_2.1_pe files. They are prompt enhancers, language models for ComfyUI’s TextGenerate node, not text encoders. Loading one as the encoder gives garbage or a shape error. The int8 files run fast only on PyTorch built for CUDA 13.0 (cu130); on cu128 they were slower than a GGUF in Kijai’s test.

ComfyUI/models
models/
├── diffusion_models/
│   ├── qwen_image_fp8_e4m3fn.safetensors
│   └── qwen_image_2.1_int8_convrot.safetensors
├── loras/
│   └── Qwen-Image-Lightning-8steps-V2.0-bf16.safetensors
├── text_encoders/
│   ├── qwen_2.5_vl_7b_fp8_scaled.safetensors
│   └── qwen3vl_8b_int8_convrot.safetensors
└── vae/
    ├── qwen_image_vae.safetensors
    └── qwen_image_2.1_vae_bf16.safetensors

Smaller files: GGUF

Qwen-Image and 2512 GGUFs load through the ComfyUI-GGUF nodes and go in diffusion_models. For the encoder, unsloth’s Qwen2.5-VL 7B GGUF is 4.7 GB at Q4_K_M and 8.1 GB at Q8. unsloth also has Qwen-Image 2.1 GGUFs (4.2 GB at Q4_K_M), but ComfyUI-GGUF’s main branch doesn’t load them yet.

ModelQ8_0Q5_K_MQ4_K_MQ4_0
Qwen-Image21.8 GB14.9 GB13.1 GB11.9 GB
251221.8 GB15.0 GB13.2 GB11.9 GB

city96 goes lower too: Q3_K_M is 9.7 GB and Q2_K 7.1 GB.

Qwen-Image-Edit

Edit is its own model for changing an existing picture. It uses the same Qwen2.5-VL 7B encoder and Qwen-Image VAE as above, plus its own file from Comfy-Org’s Edit repo: qwen_image_edit_2511_fp8mixed.safetensors is the newest, 20.5 GB. In ComfyUI the picture goes in through the TextEncodeQwenImageEditPlus node, which takes the prompt and up to three images. If you use a GGUF encoder for Edit, it needs its matching mmproj file next to it, or the image input breaks. The ComfyUI Edit tutorial has the full walkthrough. Qwen-Image 2.1 can also edit, from reference images, with the files above.

What fits your computer.

At about 1 megapixel. The text encoder runs first and ComfyUI moves it out of the way before sampling, so the model file sets the limit. A model that doesn’t fit still runs with part of it in system RAM, only slower.

  • 6 GBNo

    Even Qwen-Image Q2_K is 7.1 GB. Z-Image Turbo or Flux.2 Klein 4B are the better fit.

  • 8 GBOffloads

    Only with most of the model in system RAM. Qwen-Image 2.1 int8 (7.3 GB) comes closest. Plan on 32 GB of RAM or more.

  • 12 GBOffloads

    Qwen-Image as GGUF Q4_0 (11.9 GB) or Q3_K_M (9.7 GB), partly in RAM. Qwen-Image 2.1 int8 fits better.

  • 16 GBFits

    Qwen-Image 2.1 int8 with the int8 encoder: an RTX 4060 Ti 16 GB peaked around 15 GB with no offloading. Qwen-Image as Q4_K_M or Q5_K_M GGUF, or fp8 with offloading and the Lightning LoRA.

  • 24 GBFits

    Qwen-Image or 2512 in fp8. Comfy’s own timings are from a 24 GB RTX 4090D at 86% of its memory.

  • Mac 24 GBSwaps

    Qwen-Image 2.1 swaps hard: one edit with three references at 1280 × 736 took 35 minutes on an M4 Pro with 24 GB.

  • Mac 32 to 48 GBTight

    Qwen-Image as GGUF (13.1 to 21.8 GB) with a GGUF encoder, or Qwen-Image 2.1 in bf16, about 32 GB for the set. We found no measured Mac times to quote.

  • Mac 64 GB+Fits

    Qwen-Image 2.1 in bf16 with room to spare, or Qwen-Image as Q8_0 GGUF.

Set it up.

  1. Update ComfyUI

    Qwen-Image needs a ComfyUI from August 2025 or later; older ones say qwen_image is not in the list of CLIP types. Qwen-Image 2.1 needs ComfyUI 0.37.0 or the nightly build, for its TextEncodeQwenImage21 node. ComfyUI Desktop and the portable build can lag behind. In a manual install:

    Terminal, in the ComfyUI folder
    git pull
    pip install -r requirements.txt
  2. Download the files

    Model, text encoder and VAE from the lists above, plus a Lightning LoRA if you want 8 or 4 steps.

  3. Put them in their folders

    Model in diffusion_models, encoder in text_encoders, VAE in vae, LoRA in loras. Restart ComfyUI so the loaders list them.

  4. Open a Qwen template

    In the template browser, pick Qwen-Image: Text to Image, Qwen Image 2512 or the Qwen-Image 2.1 text to image template. The first two ship with a Lightning LoRA already wired in.

  5. Check the loaders

    Load Diffusion Model gets the model. Load CLIP gets the encoder with type qwen_image. Load VAE gets the matching VAE. For Qwen-Image and 2512 there is a ModelSamplingAuraFlow node at shift 3.1; for 2.1 the prompt goes through TextEncodeQwenImage21 instead of a plain text encode.

  6. Write a prompt and run

    Plain, detailed sentences work. The first run loads everything and takes longest.

Settings that work.

Qwen-Image and 2512

Steps
50
CFG
4
Sampler
euler
Scheduler
simple
Size
1328 × 1328
Shift
AuraFlow 3.1
Negative
yes
Latent
EmptySD3LatentImage

That is the model card: 50 steps at a true CFG of 4. Comfy’s first Qwen-Image template ran 20 steps at CFG 2.5, which is much faster and still good; its note says to try 50 for the original settings. The 2512 template uses 50 steps at CFG 4.

With a Lightning LoRA

Steps
8 or 4
CFG
1
Sampler
euler
Scheduler
simple

Steps match the LoRA’s name. Comfy’s current Qwen-Image template runs the 8-step LoRA at 8 steps, CFG 1, shift 3.1. Don’t combine Lightning with the community qwen_image_distill_full model: that one is already distilled, and runs at 15 steps and CFG 1 on its own.

AspectSize
1:11328 × 1328
16:91664 × 928
9:16928 × 1664
4:31472 × 1140
3:41140 × 1472
3:21584 × 1056
2:31056 × 1584

The model card’s sizes. Comfy’s 2512 note uses 1472 × 1104 for 4:3, a multiple of 16. Users report anything from 768 to 2048 on a side works well.

Qwen-Image 2.1

Steps
25 to 40
CFG
1
Sampler
euler
Scheduler
simple
Size
1024 to 2048
Multiple of
32
Text encode
TextEncodeQwenImage21
Latent
EmptyLatentImage

The model card uses 40 steps at 2048 × 2048. Comfy’s template uses 25 steps at 1024 × 1024, which is quicker; users who saw faint banding fixed it with 40 steps at 2K. Keep CFG at 1 unless you add a negative prompt, then raise it. For a transparent PNG, wrap the prompt the way the model card does: start with “This is an RGBA image with transparency.” and end with “The image has alpha channel and the background is transparent.”

How fast.

GPUModelStepsTime
RTX 4090D 24 GBQwen-Image fp82071 s[1]
RTX 4090D 24 GBfp8 + 8-step Lightning834 s[1]
RTX 5070 Ti 16 GB2512 GGUF Q5_K_M, 1024 px40148 s[2]
RTX 4060 Ti 16 GB2.1 int8, text to imagenot given20 s[3]
RTX 4060 Ti 16 GB2.1 int8, editnot given60 s[3]

Times after the first run, which loads the models (the 4090D took 94 s and 55 s on its first runs).

On an RTX 20-series card it is a different story: those GPUs have no bf16, fp16 overflows to black, so Qwen-Image runs in fp32 and one image took more than 10 minutes on a 2080 Ti.

When it goes wrong.

CLIPLoader ... Value not in list: type: 'qwen_image' not in [...]
ComfyUI is too old for Qwen-Image. Update it.
The image turns black partway through sampling
Start ComfyUI without --use-sage-attention, and don’t set the weight dtype to fp8_e4m3fn_fast. The same fix applies to black Lightning images.
Error while deserializing header: header too large
A broken download, usually the VAE. Download it again.
lora key not loaded, or blurry low-contrast Lightning images
Your ComfyUI is older than the LoRA support. Update it.
An error with a Qwen2.5-VL file from elsewhere
Wrong text encoder. Use Comfy-Org’s qwen_2.5_vl_7b_fp8_scaled or the full-precision one from the same repo.
TextEncodeQwenImage21 or QwenImage21Cache missing
Qwen-Image 2.1 needs ComfyUI 0.37.0 or the nightly build.
Given normalized_shape=[4096] ... got input of size [1, 338, 5120]
Qwen-Image 2.1 got the wrong encoder, often a qwen3.5_9b prompt-enhancer file. Use qwen3vl_8b_int8_convrot or qwen3vl_8b_bf16.
Qwen-Image 2.1 looks yellow, or a fine diamond grid shows on skin
For the tint, raise CFG and add a negative prompt. The grid comes from the 2.1 VAE; madebyollin’s texture-fix VAE is a drop-in replacement.

With the 4-step Lightning LoRA every seed gives nearly the same picture. Is there a way to get more variety?

Hugging Face, Qwen-Image-Lightning

The licence says non-commercial, the team says outputs are fine. Which is it for Qwen-Image 2.1?

Hugging Face, Qwen-Image-2.1

For the first, users VAE-encode a blank image and sample it at a denoise of about 0.95 instead of starting from an empty latent. The second has no official answer yet.

Questions.

Which text encoder does Qwen-Image use?

Qwen-Image and Qwen-Image 2512 use Qwen2.5-VL 7B, loaded with Load CLIP set to type qwen_image. Qwen-Image 2.1 uses Qwen3-VL 8B. The qwen3.5_9b files in the 2.1 repo are prompt enhancers, not encoders.

How much VRAM does Qwen-Image need?

24 GB runs the 20.4 GB fp8 model without offloading. 16 GB works with a Q4_K_M or Q5_K_M GGUF, or fp8 with part of it in system RAM. Qwen-Image 2.1 is smaller: its int8 files peaked around 15 GB on a 16 GB card.

Can I use Qwen-Image commercially?

Qwen-Image and Qwen-Image 2512 are Apache 2.0, so yes. Qwen-Image 2.1 is under the Qwen Research License, for non-commercial purposes only, with a separate commercial licence from Qwen.

What is the difference between Qwen-Image and 2512?

2512 is the December 2025 update of the same model, with better human realism and text. It uses the same encoder and VAE and has the same file size. Some users prefer the original for flat anime styles.

Which Lightning LoRA should I use?

The one made for your model: Qwen-Image-Lightning for the original, Qwen-Image-2512-Lightning for 2512. Run the 8-step LoRA at 8 steps and the 4-step one at 4, both at CFG 1.

Do Qwen-Image 2.1 GGUFs work in ComfyUI?

Not with the main branch of ComfyUI-GGUF yet; support is in open pull requests. Use the int8 files (7.3 GB model, 9.4 GB encoder) for now. Qwen-Image and 2512 GGUFs work.

How do I edit images with Qwen?

Either with Qwen-Image-Edit 2511, a separate model that shares Qwen-Image’s encoder and VAE, through ComfyUI’s Qwen-Image-Edit templates. Or with Qwen-Image 2.1, which edits from reference images with the same file it uses for text to image.

Sources: Qwen-Image model card, Qwen-Image 2512 model card, Qwen-Image 2.1 model card, ComfyUI Qwen-Image tutorial, ComfyUI Qwen-Image 2.1 tutorial, Comfy Qwen-Image template [1], 2512 speed thread [2], Qwen-Image 2.1 on a 4060 Ti [3], Qwen-Image-Lightning, ComfyUI issue #10852, ComfyUI issue #16470, 2.1 on an M4 Pro.

HEISS UI

Every Qwen, told apart.

HEISS UI runs Qwen-Image, Qwen-Image 2512 and Qwen-Image 2.1 on the ComfyUI you already have. Pick a file and it brings the text encoder and VAE that belong to it.

  • The right parts for each Qwen. Each version gets its own text encoder, even when the file is renamed.
  • Edit from your own pictures. Qwen-Image 2.1 takes reference images right in the composer.
  • Missing parts, shown first. Each one listed with its size and a button. Get all checks free space, and downloads resume and are verified.
  • Upscale and compare. One click to upscale with SeedVR2, then drag a slider to see what changed.

Qwen-Image Edit runs as your own workflow. Free and open source. macOS, Windows and Linux. Runs on your ComfyUI.

HEISS UI with a gallery of generated images and the prompt composer at the bottom.