Image model

How to run Flux.2 Klein locally.

Flux.2 Klein runs in ComfyUI with three files: the model, a Qwen3 text encoder and the Flux.2 VAE. The 4B version fits an 8 GB graphics card in fp8, makes an image in four steps and is Apache 2.0. The 9B is sharper, needs 12 to 16 GB and is non-commercial.

Updated 29 Sep 20269 min read

Maker
Black Forest Labs
Released
Jan 202615 January, 4B and 9B together
Licence
4B: Apache 2.09B: FLUX Non-Commercial
Memory
8 GB and up4B in fp8. 9B from 12 GB

4B or 9B.

Klein is the small end of Black Forest Labs’ Flux.2 line, released on 15 January 2026. There are two sizes, and each comes in two versions.

  • Distilled is the one to use. Four steps, CFG 1, no negative prompt. This is the file called just flux-2-klein-4b or flux-2-klein-9b.
  • Base is the undistilled model, made for training LoRAs and for more variety. It takes 20 to 50 steps at a real CFG and follows a negative prompt. The file names say base.

Pick 4B for an 8 or 12 GB card, a Mac with 16 to 24 GB, or anything you want to use commercially. Pick 9B when you have 16 GB or more and the work is personal. Both edit from reference images with the same model that makes new ones.

Files you need.

One model, one text encoder, one VAE. The text encoder has to match the size: 4B takes Qwen3 4B, 9B takes Qwen3 8B. Mixing them up is the most common Klein error.

Klein 4B

  • Model flux-2-klein-4b-fp8.safetensors ComfyUI/models/diffusion_models/ or flux-2-klein-4b.safetensors, full precision, 7.8 GB
    4.1 GB Download
  • Text encoder qwen_3_4b.safetensors ComfyUI/models/text_encoders/
    8.0 GB Download
  • VAE flux2-vae.safetensors ComfyUI/models/vae/
    0.3 GB Download

Klein 9B

  • Model flux-2-klein-9b-fp8.safetensors ComfyUI/models/diffusion_models/ gated: accept the licence on Hugging Face first. Full precision: 18.2 GB
    9.4 GB Download
  • Text encoder qwen_3_8b_fp8mixed.safetensors ComfyUI/models/text_encoders/
    8.7 GB Download
  • VAE flux2-vae.safetensors ComfyUI/models/vae/
    0.3 GB Download

The same VAE serves every Flux.2 model. Comfy’s newer templates use BFL’s small decoder instead (full_encoder_small_decoder.safetensors, 0.25 GB), which decodes about 1.4 times faster with little visible difference. Either works.

ComfyUI/models
models/
├── diffusion_models/
│   └── flux-2-klein-4b-fp8.safetensors
├── text_encoders/
│   └── qwen_3_4b.safetensors
└── vae/
    └── flux2-vae.safetensors

Smaller files: GGUF

For less memory, unsloth’s GGUF builds load through the ComfyUI-GGUF nodes. They go in diffusion_models like the others.

SizeQ8_0Q6_KQ5_K_MQ4_K_M
Klein 4B4.3 GB3.4 GB3.1 GB2.6 GB
Klein 9B10.0 GB7.9 GB7.0 GB5.9 GB

Q8_0 is close to the original. Q5_K_M and Q6_K are the usual middle ground, Q4_K_M the smallest most people keep.

What fits your computer.

Distilled Klein at 1024 × 1024. The text encoder runs first and ComfyUI moves it out of the way before sampling, so the model file sets the limit more than the total.

  • 6 GBTight

    Klein 4B as GGUF Q4_K_M, with a GGUF Qwen3 4B. Update ComfyUI first; old versions fail on Klein GGUFs.

  • 8 GBFits

    Klein 4B fp8. BFL’s own figure is about 8 GB for the 4B. One user runs 9B as GGUF Q8 on an 8 GB RTX 4060 with 40 GB of system RAM.

  • 12 GBFits

    Klein 4B at full precision, or 9B fp8 with part of it in system RAM.

  • 16 GBFits

    Klein 9B fp8 comfortably.

  • 24 GBFits

    Klein 9B at full precision. A 3090 peaked at 19.4 GB.

  • Mac 16 GBTight

    Klein 4B as GGUF. Close other large apps.

  • Mac 18 to 24 GBFits

    Klein 4B, full-precision file. fp8 saves disk here, not memory.

  • Mac 32 GB+Fits

    Klein 9B as GGUF. Q8_0 is 10 GB.

Set it up.

  1. Update ComfyUI

    Klein needs ComfyUI 0.9 or newer, from January 2026 on. Older versions fail on its files with a positional dim error. ComfyUI Desktop updates itself; the portable build has an update script. In a manual install:

    Terminal, in the ComfyUI folder
    git pull
    pip install -r requirements.txt
  2. Download the three files

    Model, text encoder and VAE from the lists above. For 9B, sign in to Hugging Face and accept the licence on the model page first.

  3. Put them in their folders

    Model in diffusion_models, text encoder in text_encoders, VAE in vae. Then restart ComfyUI so the loaders list them.

  4. Open the Klein template

    In ComfyUI’s template browser, pick Flux.2 [Klein] 4B: Text to Image or the 9B one. The edit templates add image inputs for references.

  5. Check the three loaders

    Load Diffusion Model gets the Klein file. Load CLIP gets the Qwen3 encoder with type flux2. It’s a single CLIP loader: Klein has no CLIP-L and no T5. Load VAE gets the Flux.2 VAE.

  6. Write a prompt and run

    Plain sentences work best. The first run loads everything and takes longer; the next ones are quick.

Settings that work.

Klein samples through SamplerCustomAdvanced with Flux2Scheduler, which sets the noise schedule from the step count and the image size. There is no shift to set.

Distilled (the normal files)

Steps
4
CFG
1
Sampler
euler
Scheduler
Flux2Scheduler
Size
1024 × 1024
Negative
none
Text encoder
Load CLIP, flux2
Latent
EmptyFlux2LatentImage

Base

Steps
20 to 50
CFG
4 to 5
Sampler
euler
Scheduler
Flux2Scheduler
Size
1024 × 1024
Negative
yes
Text encoder
Load CLIP, flux2
Latent
EmptyFlux2LatentImage

BFL’s Base model card uses 50 steps at guidance 4. Comfy’s Base templates use 20 steps at CFG 5, which is faster and close in quality. Any aspect ratio works: the scheduler adapts to the pixel count.

How fast.

GPUModelSizeTime
RTX 50904B Distilled, 4 steps1024 px1.2 s[1]
RTX 50904B Base, 20 steps1024 px17 s[1]
RTX 30909B Distilled, full precision1024 px24 s[2]

Times after the first run, which loads the models. The 3090 figure is from diffusers, not ComfyUI.

When it goes wrong.

mat1 and mat2 shapes cannot be multiplied (512x2560 and 7680x3072)
The 4B model got the wrong encoder or the wrong CLIP type. Use Qwen3 4B, and set Load CLIP to type flux2.
mat1 and mat2 shapes cannot be multiplied (512x12288 and 7680x3072)
Klein 4B with the Qwen3 8B encoder. The 8B one belongs to Klein 9B.
mat1 and mat2 shapes cannot be multiplied (512x4096 and 12288x4096)
Klein 9B loaded through a Flux.1-style dual CLIP loader with T5. Use one Load CLIP node, type flux2, with Qwen3 8B.
Got [32, 32, 32, 32] but expected positional dim 64
ComfyUI is too old for Klein, often with GGUF files. Update ComfyUI to 0.9 or newer.
'NoneType' object has no attribute 'Params'
An old comfy-kitchen package. Update ComfyUI’s requirements (pip install -r requirements.txt).
Overexposed images that ignore the prompt, with a GGUF encoder
A bug in ComfyUI 0.33.1 with GGUF Qwen3 encoders. Use the safetensors encoder until it’s fixed.
The negative prompt does nothing
Distilled Klein runs at CFG 1, where negatives have no effect. Use a Base file with CFG 4 to 5 if you need one.

Which text encoder goes with the 9B Q4_K_M? Matching the model and encoder files feels random.

Hugging Face, unsloth Klein 9B GGUF

Why is only the 4B under Apache 2.0, when the announcement sounded like the whole Klein family?

Hugging Face, FLUX.2-klein-9B

Questions.

Which text encoder does Flux.2 Klein use?

Klein 4B uses Qwen3 4B and Klein 9B uses Qwen3 8B, loaded with a single Load CLIP node set to type flux2. It doesn’t use Mistral like Flux.2 Dev, or CLIP-L and T5 like Flux.1.

Should I get Klein 4B or 9B?

4B on 8 to 12 GB cards, on Macs up to 24 GB, and for anything commercial, since only the 4B is Apache 2.0. 9B on 16 GB or more for personal work: it’s sharper and follows prompts a little better.

Can I use Flux.2 Klein commercially?

Klein 4B and 4B Base are Apache 2.0, so yes. Klein 9B is under the FLUX Non-Commercial License, and Black Forest Labs asks for a paid licence for client work or products you charge for.

Do Flux.1 LoRAs work on Klein?

No. Flux.2 is a new architecture, so Flux.1 LoRAs don’t load. LoRAs for Klein 4B and Klein 9B don’t cross over either, because the model sizes differ.

Why doesn’t the negative prompt work?

Distilled Klein runs at CFG 1, where a negative prompt has no effect. The Base models use a real CFG of 4 to 5, and there the negative prompt works.

Does Flux.2 Klein run on a Mac?

Yes, in ComfyUI on Apple Silicon. Use the full-precision 4B file (7.8 GB) or a GGUF: fp8 files load at full precision on a Mac, so they save disk space but not memory. 18 GB of memory is comfortable for 4B.

Sources: FLUX.2 Klein 4B model card, Klein 9B model card, BFL flux2 repository, ComfyUI Klein tutorial [1], Klein 9B memory report [2], encoder mismatch thread, ComfyUI issue #12006.

HEISS UI

Klein, without the wiring.

HEISS UI runs Flux.2 Klein on the ComfyUI you already have. Pick the file, write a prompt, and it handles the rest.

  • No wrong-encoder errors. Drop in any Klein file and it gets the text encoder that fits it.
  • A first model that fits. On an empty studio it’s one tap away, with the version for your GPU or Mac marked.
  • Edit from your own pictures. Add reference images in the composer and describe the change.
  • Missing parts, shown first. Each one listed with its size and a button. Get all checks free space, and downloads resume and are verified.

Free and open source. macOS, Windows and Linux. Runs on your ComfyUI.

HEISS UI with a gallery of generated images and the prompt composer at the bottom.