Dev, Schnell or Krea Dev.
Flux.1 is Black Forest Labs’ first open model family, released on 1 August 2024: a 12B image model with CLIP-L and T5-XXL as text encoders. Everything below shares that architecture, so the encoders and the VAE are the same for all of them.
- Dev is the quality model. It’s guidance-distilled: it runs at a guidance value of about 3.5, CFG 1 and no negative prompt, in 20 to 50 steps.
- Schnell is the fast one. Four steps, CFG 1, and Apache 2.0, so you can use it commercially.
- Krea Dev is a Dev fine-tune by BFL and Krea from July 2025, made to look less like AI. It uses Dev’s settings and takes Flux.1 LoRAs, with mixed results.
- De-distilled fine-tunes are community models that bring back a real CFG and a working negative prompt, at the cost of more steps.
For a newer model at the same memory, look at Flux.2 Klein. Flux.1 LoRAs don’t carry over to Flux.2, though, and the LoRA library for Flux.1 is large.
Files you need.
A model, two text encoders and a VAE. The encoders go into one DualCLIPLoader node with type flux.
-
Model23.8 GB Download
flux1-dev.safetensorsComfyUI/models/diffusion_models/ -
Text encoder0.2 GB Download
clip_l.safetensorsComfyUI/models/text_encoders/ -
Text encoder5.2 GB Download
t5xxl_fp8_e4m3fn_scaled.safetensorsComfyUI/models/text_encoders/ or t5xxl_fp16.safetensors, 9.8 GB, with more than 32 GB of RAM -
VAE0.3 GB Download
ae.safetensorsComfyUI/models/vae/
Other versions
-
Schnell23.8 GB Download
flux1-schnell.safetensorsComfyUI/models/diffusion_models/ -
Krea Dev11.9 GB Download
flux1-krea-dev_fp8_scaled.safetensorsComfyUI/models/diffusion_models/ full precision: 23.8 GB, gated at BFL -
All in one17.3 GB Download
flux1-dev-fp8.safetensorsComfyUI/models/checkpoints/ Dev in fp8 with encoders and VAE inside. Schnell: flux1-schnell-fp8, 17.2 GB
The all-in-one checkpoints are the simplest start: one file in checkpoints, loaded with Load Checkpoint. The separate files give you more control over which T5 you use.
models/
├── diffusion_models/
│ └── flux1-dev.safetensors
├── text_encoders/
│ ├── clip_l.safetensors
│ └── t5xxl_fp8_e4m3fn_scaled.safetensors
└── vae/
└── ae.safetensors
Smaller files: GGUF
city96’s GGUF builds load through the ComfyUI-GGUF nodes and go in diffusion_models. There’s no Q4_K_M or Q5_K_M for Flux.1 Dev.
| Quant | Size | Quant | Size |
|---|---|---|---|
| Q8_0 | 12.7 GB | Q4_K_S | 6.8 GB |
| Q6_K | 9.9 GB | Q3_K_S | 5.2 GB |
| Q5_K_S | 8.3 GB | Q2_K | 4.0 GB |
Q4_0 is the same size as Q4_K_S. Q8_0 is close to the original. Q5_K_S loses very little, Q3 still holds up, and Q2 is where it breaks. Keep the T5 at fp8 rather than a low-bit T5 GGUF, which hurts prompt understanding.
What fits your computer.
Flux.1 Dev at 1024 × 1024 with the fp8 T5. ComfyUI runs the encoders first and moves them aside before sampling, so the model file sets the limit.
- 6 GBTight
Q3_K_S or Q4 GGUF with the fp8 T5, and 32 GB of system RAM.
- 8 GBFits
Q4_0 GGUF with fp8 T5: an RTX 2080 used 6.4 GB of graphics memory and needed 32 GB of RAM. Schnell ran in under a minute on an RTX 4060.
- 12 GBFits
Q8_0 GGUF or the fp8 files. The full model also runs with part of it in system RAM.
- 16 GBFits
fp8 comfortably, including the fp8 Krea Dev and the all-in-one checkpoint.
- 24 GBFits
Full precision with the weight type set to
fp8_e4m3fnin Load Diffusion Model. At the default type, model and fp16 T5 overflow 24 GB and it slows to minutes. - Mac 16 GBTight
Schnell as a Q5 GGUF runs on an M1 with 16 GB.
- Mac 24 GBSlow
Q4 GGUF runs. An M2 Air took about 52 seconds per step.
- Mac 32 GB+Fits
The full-precision files or Q8_0 GGUF. fp8 files don’t work on a Mac.
Set it up.
-
Download the files
Model, CLIP-L, T5 and VAE from the list above, or one all-in-one checkpoint. For the gated BFL originals, sign in to Hugging Face and accept the licence first; the Comfy-Org copies linked here aren’t gated.
-
Put them in their folders
Model in
diffusion_models, both encoders intext_encoders, VAE invae. An all-in-one checkpoint goes incheckpoints. Restart ComfyUI so the loaders list them. -
Open a Flux.1 template
In ComfyUI’s template browser, pick Flux.1 Dev: Text to Image, Flux.1 Schnell Full: Text to Image or Flux.1 Krea Dev. For the all-in-one files, Flux.1 Schnell FP8 or Flux.1 Dev fp8: Text to Image.
-
Check the loaders
Load Diffusion Model gets the model; on 24 GB or less, set its weight type to
fp8_e4m3fn. DualCLIPLoader gets CLIP-L and T5 with typeflux. Load VAE getsae.safetensors. For a GGUF, swap the model loader for Unet Loader (GGUF). -
Keep CFG at 1
Dev and Schnell have no negative prompt. The KSampler’s CFG stays at 1; the prompt strength is the guidance value, 3.5 by default.
-
Write a prompt and run
Plain descriptive sentences work well. T5 reads up to 512 tokens on Dev and 256 on Schnell, so long prompts are fine.
Settings that work.
Dev and Krea Dev
- Steps
- 20 to 50
- Guidance
- 3.5
- CFG
- 1
- Sampler
- euler
- Scheduler
- simple
- Size
- 1024 × 1024
- Negative
- none
- T5 length
- 512 tokens
BFL’s model card uses 50 steps at guidance 3.5. Comfy’s templates use 20 steps and have no FluxGuidance node, so ComfyUI’s default of 3.5 applies. Add a FluxGuidance node to change it. The latent is EmptySD3LatentImage. Stay near 1 megapixel: beyond about 2 megapixels, images degrade.
Schnell
- Steps
- 4
- CFG
- 1
- Sampler
- euler
- Scheduler
- simple
Guidance and CFG are different things. Guidance is a value baked into Dev’s training that you pass in as conditioning. CFG is the sampler’s classifier-free guidance, and on Dev and Schnell it must stay at 1, or images turn blurry and burnt. For a working negative prompt, use a de-distilled fine-tune at a real CFG and follow its page for steps.
How fast.
| GPU | Setup | Time |
|---|---|---|
| RTX 4090 | Dev, fp8_e4m3fn, 1024 px | 14 s per image[1] |
| RTX 4070 12 GB | Dev, default type | 2.5 s per step[1] |
| RTX 3060 12 GB | Dev, default type | 5 s per step[1] |
| RTX 4070 Super | Q4_0 GGUF | 1.9 s per step[2] |
| RTX 2080 8 GB | Q4_0 GGUF, fp8 T5 | 3.2 s per step[2] |
| Mac Studio 96 GB | Q4_K_S, 12 steps | 355 s[2] |
User reports. The 3060 figure is with NVIDIA’s system memory fallback turned off. On the 4090, the default weight type took up to 10 minutes, because the model spilled out of graphics memory.
When it goes wrong.
mat1 and mat2 shapes cannot be multiplied (1x1280 and 768x3072)- The DualCLIPLoader type is set to
sdxl. Set it toflux. - Blurry, washed-out or burnt images
- CFG above 1, often with a negative prompt, on a distilled model. Set CFG to 1 and leave the negative empty.
- Very slow on a 24 GB card
- The full-precision model and the fp16 T5 don’t fit together. Set the weight type to
fp8_e4m3fnin Load Diffusion Model, or use the fp8 T5. Trying to convert Float8_e4m3fn to the MPS backend- An fp8 file on a Mac. Use the full-precision files or a GGUF.
- A black image
- The VAE decoded in float16. Run the VAE in bf16 or fp32.
clip missing: ['text_projection.weight']- A harmless warning. Nothing to fix.
Token indices sequence length is longer than the specified maximum sequence length- CLIP-L stops at 77 tokens and says so. T5 still reads the whole prompt, so ignore it.
- A LoRA looks weaker with a GGUF model
- In low-VRAM mode, part of the LoRA isn’t applied with GGUF models. Give ComfyUI more room or use an fp8 model.
My Flux images in Forge come out blurry. What am I doing wrong? The answer: set CFG to 1 and don’t use a negative prompt.
Is Q4 good enough, or do I need Q8? Where does the quality actually start to break?
Questions.
What’s the difference between Flux.1 Dev and Schnell?
Dev is guidance-distilled and makes better images in 20 to 50 steps. Schnell is timestep-distilled for 1 to 4 steps, so it’s much faster, and it’s Apache 2.0. Both use the same encoders and VAE.
Can I use Flux.1 Dev commercially?
Schnell, yes: it’s Apache 2.0. Dev and Krea Dev are under the FLUX [dev] Non-Commercial License. The model card says outputs may be used commercially, but the licence text conflicts with that, so ask Black Forest Labs before running Dev for paid client work.
Why doesn’t the negative prompt work in Flux?
Dev and Schnell are distilled and run at CFG 1, where the negative prompt has no effect. Raising CFG makes images blurry. De-distilled fine-tunes restore a real CFG and a working negative prompt.
Should I use T5 fp16 or fp8?
The fp8 T5 (5.2 GB) saves memory with little loss. ComfyUI’s own examples recommend the fp16 one (9.8 GB) if you have more than 32 GB of system RAM. Avoid low-bit T5 GGUFs: they noticeably hurt prompt understanding.
Which Flux.1 GGUF should I pick?
Q8_0 (12.7 GB) for quality close to the original, Q6_K or Q5_K_S as the middle ground, and Q4_0 or Q4_K_S for an 8 GB card. Quality holds up to Q3 and breaks at Q2.
Does Flux.1 run on a Mac?
Yes, in ComfyUI on Apple Silicon, with the full-precision files or a GGUF. fp8 files don’t work on a Mac. Expect it to be slow: an M1 Max took about two minutes for a four-step Schnell image.
Do Flux.1 LoRAs work on Krea Dev or Flux.2?
On Krea Dev they load, since it has the same architecture, with mixed results. On Flux.2 they don’t load at all: it’s a new architecture.
Sources: FLUX.1 Dev model card, FLUX.1 Schnell model card, FLUX.1 Krea Dev model card, ComfyUI Flux examples, FLUX.1-dev speed thread [1], Flux GGUF speed thread [2], licence thread, macOS thread, ComfyUI issue #6969.