Image model

How to run Z-Image locally.

Z-Image Turbo runs in ComfyUI with three files: the 12.3 GB model, the Qwen3 4B text encoder and the Flux.1 VAE. It makes an image in eight steps, runs at full precision on a 16 GB card, and on 8 GB as int8 or GGUF. Turbo and Base are both Apache 2.0.

Updated 29 Sep 20268 min read

Maker
Tongyi-MAI, Alibaba
Released
Nov 2025Turbo in November, Base in January 2026
Licence
Apache 2.0Turbo and Base, commercial use allowed
Memory
8 GB and up16 GB at full precision

Turbo or Base.

Z-Image is a 6B single-stream model from Alibaba’s Tongyi lab. There are two versions, and they share the text encoder and VAE.

  • Turbo, from late November 2025, is the one most people use. Eight steps, CFG 1, no negative prompt. The files say turbo.
  • Base, from January 2026, is the undistilled model. Officially it’s just “Z-Image”; ComfyUI’s file is z_image_bf16. It takes 30 to 50 steps at CFG 3 to 5, follows a negative prompt, and gives more variety. It’s also the one to train LoRAs on.

An edit model, Z-Image-Edit, was announced alongside them but isn’t on Hugging Face yet.

Files you need.

One model, one text encoder, one VAE. The encoder must be Qwen3 4B, or a fine-tune of it. A smaller Qwen3 doesn’t work.

Turbo

  • Model z_image_turbo_bf16.safetensors ComfyUI/models/diffusion_models/ smaller: z_image_turbo_int8_convrot.safetensors, 6.2 GB
    12.3 GB Download
  • Text encoder qwen_3_4b.safetensors ComfyUI/models/text_encoders/ smaller: qwen_3_4b_fp8_mixed.safetensors, 5.6 GB
    8.0 GB Download
  • VAE ae.safetensors ComfyUI/models/vae/
    0.3 GB Download

Base

  • Model z_image_bf16.safetensors ComfyUI/models/diffusion_models/ smaller: z_image_int8_convrot.safetensors, 6.2 GB. Same encoder and VAE as Turbo
    12.3 GB Download

The VAE is Flux.1’s ae.safetensors; if you already run Flux.1 or Chroma, you have it. The Qwen3 4B encoder is the same file Flux.2 Klein 4B uses. Get Comfy-Org’s converted files rather than the diffusers files from Tongyi’s own repo, which look broken when loaded the wrong way. The int8 files run fast only on PyTorch built for CUDA 13.0 (cu130).

ComfyUI/models
models/
├── diffusion_models/
│   └── z_image_turbo_bf16.safetensors
├── text_encoders/
│   └── qwen_3_4b.safetensors
└── vae/
    └── ae.safetensors

Smaller files: GGUF

unsloth’s GGUF builds load through the ComfyUI-GGUF nodes, which treat Z-Image like Lumina 2. They go in diffusion_models. Qwen3 4B also comes as GGUF from unsloth: 2.5 GB at Q4_K_M, 4.3 GB at Q8_0.

VersionQ8_0Q6_KQ5_K_MQ4_K_M
Turbo7.2 GB5.9 GB5.6 GB5.0 GB
Base7.2 GB6.1 GB5.6 GB5.1 GB

Q8_0 is close to the original. Q5_K_M and Q6_K are the usual middle ground. Base goes down to 4.0 GB at Q2_K.

What fits your computer.

Turbo at 1024 × 1024. The text encoder runs first and ComfyUI moves it out of the way before sampling, so the model file sets the limit. The makers say Turbo “fits comfortably within 16G VRAM”.

  • 6 GBTight

    Turbo as GGUF Q4_K_M (5.0 GB) with a GGUF Qwen3 4B, and 32 GB of system RAM. On an RTX 20-series card only Turbo works: Base gives noise there.

  • 8 GBFits

    Turbo int8 (6.2 GB) or a Q5_K_M to Q6_K GGUF, with the fp8 encoder.

  • 12 GBFits

    int8 or GGUF Q8_0 with room to spare. The full 12.3 GB file runs with a little of it in system RAM.

  • 16 GBFits

    Turbo or Base at full precision.

  • 24 GBFits

    Everything, with room for LoRAs and larger sizes.

  • Mac 16 to 24 GBTight

    Turbo as GGUF, with a GGUF Qwen3 4B.

  • Mac 32 GB+Fits

    Turbo or Base at full precision: 12.3 GB model plus the 8.0 GB encoder.

Set it up.

  1. Update ComfyUI

    Old versions fail on the encoder with size mismatch for model.embed_tokens.weight. ComfyUI Desktop updates itself; the portable build has an update script. In a manual install:

    Terminal, in the ComfyUI folder
    git pull
    pip install -r requirements.txt
  2. Download the three files

    Model, text encoder and VAE from the lists above.

  3. Put them in their folders

    Model in diffusion_models, encoder in text_encoders, VAE in vae. Restart ComfyUI so the loaders list them.

  4. Open the Z-Image template

    In the template browser, pick Z-Image-Turbo: Text to Image, or Z-Image: Text to Image for Base. ComfyUI’s getting-started Text to Image template uses Z-Image Turbo too.

  5. Check the loaders

    Load Diffusion Model gets the Z-Image file. Load CLIP gets qwen_3_4b with type lumina2. Load VAE gets ae. The template adds ModelSamplingAuraFlow at shift 3, which is also ComfyUI’s default for Z-Image.

  6. Write a prompt and run

    Plain, detailed sentences work best. The first run loads everything and takes longer.

Settings that work.

Turbo

Steps
8
CFG
1
Sampler
res_multistep
Scheduler
simple
Size
1024 × 1024
Shift
AuraFlow 3
Negative
none
Latent
EmptySD3LatentImage

The model card says 9 scheduler steps with guidance 0, which comes to 8 model passes. Comfy’s 8 steps at CFG 1 is the same thing. At CFG 1 a negative prompt has no effect.

Base

Steps
30 to 50
CFG
3 to 5
Sampler
res_multistep
Scheduler
simple
Size
512 to 2048
Shift
AuraFlow 3
Negative
yes
Latent
EmptySD3LatentImage

The Base model card’s example is 50 steps at CFG 4. Comfy’s template uses 25 steps at CFG 4 and notes 30 to 50 steps and CFG 3 to 5. Below CFG 3, Base doesn’t render properly. Any aspect works between 512 and 2048 on a side. Users on the model page like CFG 4 to 4.5.

Z-Image or Flux.2 Klein.

These two get compared more than any other pair. Both are small, fast and Apache 2.0 (Klein only in its 4B size), and they use the same Qwen3 4B encoder file, so trying both costs one extra model download.

Z-Image TurboKlein 4B
Size6B4B
Steps84
Model file12.3 GB, int8 6.2 GB7.8 GB, fp8 4.1 GB
EncoderQwen3 4B, type lumina2Qwen3 4B, type flux2
VAEFlux.1 aeFlux.2 VAE
Editingnot yetreference images
LicenceApache 2.0Apache 2.0

Klein 9B is sharper, under a non-commercial licence.

Pick Klein 4B for the smallest files, four steps and editing from reference images. Pick Z-Image Turbo when you have 12 GB or more and want the larger model. Our Flux.2 Klein guide has its files and settings.

When it goes wrong.

Black images with the stock workflow
Start ComfyUI without --use-sage-attention and without Triton tricks, and don’t use the run_nvidia_gpu_fast_fp16_accumulation.bat launcher. Update ComfyUI too.
size mismatch for model.embed_tokens.weight
ComfyUI is too old for the Qwen3 4B encoder. Update with git pull and pip install -r requirements.txt.
unet missing: ['norm_final.weight'] in the console
Harmless. The images come out fine.
Base gives noise on an RTX 20-series or older AMD card
Those GPUs have no bf16. ComfyUI’s fp16 workaround works for Turbo but not for Base, so use Turbo.
mat1 and mat2 shapes cannot be multiplied (512x2560 and 12288x4096)
The model and encoder don’t belong together. In the reported case a Flux.2 Klein GGUF was loaded instead of Z-Image. Check the file in each loader.
The negative prompt does nothing
Turbo runs at CFG 1, where negatives have no effect. Use Base at CFG 3 to 5 if you need one.

Can I save memory with Qwen3 2B as the text encoder?

Hugging Face, Comfy-Org z_image_turbo

Should LoRAs be trained on Base and used on Turbo, and do they carry over at all?

Hugging Face, Tongyi-MAI Z-Image

The answer to the first is no: it has to be Qwen3 4B or a fine-tune of it. The second is still open, with results both ways.

Questions.

What is the difference between Z-Image Turbo and Base?

Turbo is distilled: 8 steps at CFG 1, no negative prompt. Base is the undistilled model: 30 to 50 steps at CFG 3 to 5, with a negative prompt and more variety. Base is also the one to train LoRAs on.

How much VRAM does Z-Image Turbo need?

The 12.3 GB full-precision file runs comfortably on 16 GB. On 8 GB use the 6.2 GB int8 file or a GGUF, and on 6 GB a Q4_K_M GGUF with 32 GB of system RAM.

Which text encoder does Z-Image use?

Qwen3 4B, loaded with Load CLIP set to type lumina2. It’s the same file Flux.2 Klein 4B uses. A smaller Qwen3 doesn’t work; fine-tunes of Qwen3 4B do.

Why does Z-Image make black images?

Most often SageAttention, Triton tweaks or the fast fp16 accumulation launcher. Start ComfyUI without them and update it.

Is Z-Image better than Flux.2 Klein?

They suit different setups. Klein 4B is smaller, runs in 4 steps and edits from reference images. Z-Image Turbo is the larger 6B model at 8 steps. Both are Apache 2.0 and share the Qwen3 4B encoder, so trying both is easy.

Can I use Z-Image commercially?

Yes. Z-Image Turbo and Z-Image Base are both Apache 2.0.

Does Z-Image run on a Mac?

Yes, in ComfyUI on Apple Silicon. Use the full-precision or GGUF files, not fp8 or int8, and update ComfyUI if it fails to load.

Sources: Z-Image-Turbo model card, Z-Image model card, ComfyUI Z-Image Turbo tutorial, Comfy-Org Z-Image Turbo files, black image thread, Base troubleshooting thread, ComfyUI issue #12176, ComfyUI issue #12132.

HEISS UI

Fast pictures, no setup.

HEISS UI runs Z-Image Turbo and Base on the ComfyUI you already have. Pick the file and it’s ready to generate.

  • Turbo or Base, spotted for you. Each gets settings that work, so there’s nothing to look up.
  • One download, two models. Z-Image shares its text encoder with Flux.2 Klein, so it only downloads once.
  • Missing parts, shown first. Each one listed with its size and a button. Get all checks free space, and downloads resume and are verified.
  • A phone studio. Prompt, browse and share from the couch while the computer renders.

Free and open source. macOS, Windows and Linux. Runs on your ComfyUI.

HEISS UI with a gallery of generated images and the prompt composer at the bottom.