4B or 9B.
Klein is the small end of Black Forest Labs’ Flux.2 line, released on 15 January 2026. There are two sizes, and each comes in two versions.
- Distilled is the one to use. Four steps, CFG 1, no negative prompt. This is the file called just
flux-2-klein-4borflux-2-klein-9b. - Base is the undistilled model, made for training LoRAs and for more variety. It takes 20 to 50 steps at a real CFG and follows a negative prompt. The file names say
base.
Pick 4B for an 8 or 12 GB card, a Mac with 16 to 24 GB, or anything you want to use commercially. Pick 9B when you have 16 GB or more and the work is personal. Both edit from reference images with the same model that makes new ones.
Files you need.
One model, one text encoder, one VAE. The text encoder has to match the size: 4B takes Qwen3 4B, 9B takes Qwen3 8B. Mixing them up is the most common Klein error.
Klein 4B
-
Model4.1 GB Download
flux-2-klein-4b-fp8.safetensorsComfyUI/models/diffusion_models/ or flux-2-klein-4b.safetensors, full precision, 7.8 GB -
Text encoder8.0 GB Download
qwen_3_4b.safetensorsComfyUI/models/text_encoders/ -
VAE0.3 GB Download
flux2-vae.safetensorsComfyUI/models/vae/
Klein 9B
-
Model9.4 GB Download
flux-2-klein-9b-fp8.safetensorsComfyUI/models/diffusion_models/ gated: accept the licence on Hugging Face first. Full precision: 18.2 GB -
Text encoder8.7 GB Download
qwen_3_8b_fp8mixed.safetensorsComfyUI/models/text_encoders/ -
VAE0.3 GB Download
flux2-vae.safetensorsComfyUI/models/vae/
The same VAE serves every Flux.2 model. Comfy’s newer templates use BFL’s small decoder instead (full_encoder_small_decoder.safetensors, 0.25 GB), which decodes about 1.4 times faster with little visible difference. Either works.
models/
├── diffusion_models/
│ └── flux-2-klein-4b-fp8.safetensors
├── text_encoders/
│ └── qwen_3_4b.safetensors
└── vae/
└── flux2-vae.safetensors
Smaller files: GGUF
For less memory, unsloth’s GGUF builds load through the ComfyUI-GGUF nodes. They go in diffusion_models like the others.
| Size | Q8_0 | Q6_K | Q5_K_M | Q4_K_M |
|---|---|---|---|---|
| Klein 4B | 4.3 GB | 3.4 GB | 3.1 GB | 2.6 GB |
| Klein 9B | 10.0 GB | 7.9 GB | 7.0 GB | 5.9 GB |
Q8_0 is close to the original. Q5_K_M and Q6_K are the usual middle ground, Q4_K_M the smallest most people keep.
What fits your computer.
Distilled Klein at 1024 × 1024. The text encoder runs first and ComfyUI moves it out of the way before sampling, so the model file sets the limit more than the total.
- 6 GBTight
Klein 4B as GGUF Q4_K_M, with a GGUF Qwen3 4B. Update ComfyUI first; old versions fail on Klein GGUFs.
- 8 GBFits
Klein 4B fp8. BFL’s own figure is about 8 GB for the 4B. One user runs 9B as GGUF Q8 on an 8 GB RTX 4060 with 40 GB of system RAM.
- 12 GBFits
Klein 4B at full precision, or 9B fp8 with part of it in system RAM.
- 16 GBFits
Klein 9B fp8 comfortably.
- 24 GBFits
Klein 9B at full precision. A 3090 peaked at 19.4 GB.
- Mac 16 GBTight
Klein 4B as GGUF. Close other large apps.
- Mac 18 to 24 GBFits
Klein 4B, full-precision file. fp8 saves disk here, not memory.
- Mac 32 GB+Fits
Klein 9B as GGUF. Q8_0 is 10 GB.
Set it up.
-
Update ComfyUI
Klein needs ComfyUI 0.9 or newer, from January 2026 on. Older versions fail on its files with a positional dim error. ComfyUI Desktop updates itself; the portable build has an update script. In a manual install:
Terminal, in the ComfyUI foldergit pull pip install -r requirements.txt
-
Download the three files
Model, text encoder and VAE from the lists above. For 9B, sign in to Hugging Face and accept the licence on the model page first.
-
Put them in their folders
Model in
diffusion_models, text encoder intext_encoders, VAE invae. Then restart ComfyUI so the loaders list them. -
Open the Klein template
In ComfyUI’s template browser, pick Flux.2 [Klein] 4B: Text to Image or the 9B one. The edit templates add image inputs for references.
-
Check the three loaders
Load Diffusion Model gets the Klein file. Load CLIP gets the Qwen3 encoder with type
flux2. It’s a single CLIP loader: Klein has no CLIP-L and no T5. Load VAE gets the Flux.2 VAE. -
Write a prompt and run
Plain sentences work best. The first run loads everything and takes longer; the next ones are quick.
Settings that work.
Klein samples through SamplerCustomAdvanced with Flux2Scheduler, which sets the noise schedule from the step count and the image size. There is no shift to set.
Distilled (the normal files)
- Steps
- 4
- CFG
- 1
- Sampler
- euler
- Scheduler
- Flux2Scheduler
- Size
- 1024 × 1024
- Negative
- none
- Text encoder
- Load CLIP, flux2
- Latent
- EmptyFlux2LatentImage
Base
- Steps
- 20 to 50
- CFG
- 4 to 5
- Sampler
- euler
- Scheduler
- Flux2Scheduler
- Size
- 1024 × 1024
- Negative
- yes
- Text encoder
- Load CLIP, flux2
- Latent
- EmptyFlux2LatentImage
BFL’s Base model card uses 50 steps at guidance 4. Comfy’s Base templates use 20 steps at CFG 5, which is faster and close in quality. Any aspect ratio works: the scheduler adapts to the pixel count.
How fast.
| GPU | Model | Size | Time |
|---|---|---|---|
| RTX 5090 | 4B Distilled, 4 steps | 1024 px | 1.2 s[1] |
| RTX 5090 | 4B Base, 20 steps | 1024 px | 17 s[1] |
| RTX 3090 | 9B Distilled, full precision | 1024 px | 24 s[2] |
Times after the first run, which loads the models. The 3090 figure is from diffusers, not ComfyUI.
When it goes wrong.
mat1 and mat2 shapes cannot be multiplied (512x2560 and 7680x3072)- The 4B model got the wrong encoder or the wrong CLIP type. Use Qwen3 4B, and set Load CLIP to type
flux2. mat1 and mat2 shapes cannot be multiplied (512x12288 and 7680x3072)- Klein 4B with the Qwen3 8B encoder. The 8B one belongs to Klein 9B.
mat1 and mat2 shapes cannot be multiplied (512x4096 and 12288x4096)- Klein 9B loaded through a Flux.1-style dual CLIP loader with T5. Use one Load CLIP node, type
flux2, with Qwen3 8B. Got [32, 32, 32, 32] but expected positional dim 64- ComfyUI is too old for Klein, often with GGUF files. Update ComfyUI to 0.9 or newer.
'NoneType' object has no attribute 'Params'- An old
comfy-kitchenpackage. Update ComfyUI’s requirements (pip install -r requirements.txt). - Overexposed images that ignore the prompt, with a GGUF encoder
- A bug in ComfyUI 0.33.1 with GGUF Qwen3 encoders. Use the safetensors encoder until it’s fixed.
- The negative prompt does nothing
- Distilled Klein runs at CFG 1, where negatives have no effect. Use a Base file with CFG 4 to 5 if you need one.
Which text encoder goes with the 9B Q4_K_M? Matching the model and encoder files feels random.
Why is only the 4B under Apache 2.0, when the announcement sounded like the whole Klein family?
Questions.
Which text encoder does Flux.2 Klein use?
Klein 4B uses Qwen3 4B and Klein 9B uses Qwen3 8B, loaded with a single Load CLIP node set to type flux2. It doesn’t use Mistral like Flux.2 Dev, or CLIP-L and T5 like Flux.1.
Should I get Klein 4B or 9B?
4B on 8 to 12 GB cards, on Macs up to 24 GB, and for anything commercial, since only the 4B is Apache 2.0. 9B on 16 GB or more for personal work: it’s sharper and follows prompts a little better.
Can I use Flux.2 Klein commercially?
Klein 4B and 4B Base are Apache 2.0, so yes. Klein 9B is under the FLUX Non-Commercial License, and Black Forest Labs asks for a paid licence for client work or products you charge for.
Do Flux.1 LoRAs work on Klein?
No. Flux.2 is a new architecture, so Flux.1 LoRAs don’t load. LoRAs for Klein 4B and Klein 9B don’t cross over either, because the model sizes differ.
Why doesn’t the negative prompt work?
Distilled Klein runs at CFG 1, where a negative prompt has no effect. The Base models use a real CFG of 4 to 5, and there the negative prompt works.
Does Flux.2 Klein run on a Mac?
Yes, in ComfyUI on Apple Silicon. Use the full-precision 4B file (7.8 GB) or a GGUF: fp8 files load at full precision on a Mac, so they save disk space but not memory. 18 GB of memory is comfortable for 4B.
Sources: FLUX.2 Klein 4B model card, Klein 9B model card, BFL flux2 repository, ComfyUI Klein tutorial [1], Klein 9B memory report [2], encoder mismatch thread, ComfyUI issue #12006.