Image model

How to run Flux.1 locally.

Flux.1 runs in ComfyUI with four files: the model, two text encoders (CLIP-L and T5-XXL) and the Flux VAE. As a GGUF it runs on an 8 GB graphics card, and the full model wants 24 GB. Schnell is Apache 2.0; Dev and Krea Dev are non-commercial.

Updated 29 Sep 202610 min read

Maker
Black Forest Labs
Released
Aug 2024Krea Dev in July 2025
Licence
Schnell: Apache 2.0Dev, Krea Dev: FLUX Non-Commercial
Memory
8 GB and upas GGUF. Full precision 24 GB

Dev, Schnell or Krea Dev.

Flux.1 is Black Forest Labs’ first open model family, released on 1 August 2024: a 12B image model with CLIP-L and T5-XXL as text encoders. Everything below shares that architecture, so the encoders and the VAE are the same for all of them.

  • Dev is the quality model. It’s guidance-distilled: it runs at a guidance value of about 3.5, CFG 1 and no negative prompt, in 20 to 50 steps.
  • Schnell is the fast one. Four steps, CFG 1, and Apache 2.0, so you can use it commercially.
  • Krea Dev is a Dev fine-tune by BFL and Krea from July 2025, made to look less like AI. It uses Dev’s settings and takes Flux.1 LoRAs, with mixed results.
  • De-distilled fine-tunes are community models that bring back a real CFG and a working negative prompt, at the cost of more steps.

For a newer model at the same memory, look at Flux.2 Klein. Flux.1 LoRAs don’t carry over to Flux.2, though, and the LoRA library for Flux.1 is large.

Files you need.

A model, two text encoders and a VAE. The encoders go into one DualCLIPLoader node with type flux.

  • Model flux1-dev.safetensors ComfyUI/models/diffusion_models/
    23.8 GB Download
  • Text encoder clip_l.safetensors ComfyUI/models/text_encoders/
    0.2 GB Download
  • Text encoder t5xxl_fp8_e4m3fn_scaled.safetensors ComfyUI/models/text_encoders/ or t5xxl_fp16.safetensors, 9.8 GB, with more than 32 GB of RAM
    5.2 GB Download
  • VAE ae.safetensors ComfyUI/models/vae/
    0.3 GB Download

Other versions

  • Schnell flux1-schnell.safetensors ComfyUI/models/diffusion_models/
    23.8 GB Download
  • Krea Dev flux1-krea-dev_fp8_scaled.safetensors ComfyUI/models/diffusion_models/ full precision: 23.8 GB, gated at BFL
    11.9 GB Download
  • All in one flux1-dev-fp8.safetensors ComfyUI/models/checkpoints/ Dev in fp8 with encoders and VAE inside. Schnell: flux1-schnell-fp8, 17.2 GB
    17.3 GB Download

The all-in-one checkpoints are the simplest start: one file in checkpoints, loaded with Load Checkpoint. The separate files give you more control over which T5 you use.

ComfyUI/models
models/
├── diffusion_models/
│   └── flux1-dev.safetensors
├── text_encoders/
│   ├── clip_l.safetensors
│   └── t5xxl_fp8_e4m3fn_scaled.safetensors
└── vae/
    └── ae.safetensors

Smaller files: GGUF

city96’s GGUF builds load through the ComfyUI-GGUF nodes and go in diffusion_models. There’s no Q4_K_M or Q5_K_M for Flux.1 Dev.

QuantSizeQuantSize
Q8_012.7 GBQ4_K_S6.8 GB
Q6_K9.9 GBQ3_K_S5.2 GB
Q5_K_S8.3 GBQ2_K4.0 GB

Q4_0 is the same size as Q4_K_S. Q8_0 is close to the original. Q5_K_S loses very little, Q3 still holds up, and Q2 is where it breaks. Keep the T5 at fp8 rather than a low-bit T5 GGUF, which hurts prompt understanding.

What fits your computer.

Flux.1 Dev at 1024 × 1024 with the fp8 T5. ComfyUI runs the encoders first and moves them aside before sampling, so the model file sets the limit.

  • 6 GBTight

    Q3_K_S or Q4 GGUF with the fp8 T5, and 32 GB of system RAM.

  • 8 GBFits

    Q4_0 GGUF with fp8 T5: an RTX 2080 used 6.4 GB of graphics memory and needed 32 GB of RAM. Schnell ran in under a minute on an RTX 4060.

  • 12 GBFits

    Q8_0 GGUF or the fp8 files. The full model also runs with part of it in system RAM.

  • 16 GBFits

    fp8 comfortably, including the fp8 Krea Dev and the all-in-one checkpoint.

  • 24 GBFits

    Full precision with the weight type set to fp8_e4m3fn in Load Diffusion Model. At the default type, model and fp16 T5 overflow 24 GB and it slows to minutes.

  • Mac 16 GBTight

    Schnell as a Q5 GGUF runs on an M1 with 16 GB.

  • Mac 24 GBSlow

    Q4 GGUF runs. An M2 Air took about 52 seconds per step.

  • Mac 32 GB+Fits

    The full-precision files or Q8_0 GGUF. fp8 files don’t work on a Mac.

Set it up.

  1. Download the files

    Model, CLIP-L, T5 and VAE from the list above, or one all-in-one checkpoint. For the gated BFL originals, sign in to Hugging Face and accept the licence first; the Comfy-Org copies linked here aren’t gated.

  2. Put them in their folders

    Model in diffusion_models, both encoders in text_encoders, VAE in vae. An all-in-one checkpoint goes in checkpoints. Restart ComfyUI so the loaders list them.

  3. Open a Flux.1 template

    In ComfyUI’s template browser, pick Flux.1 Dev: Text to Image, Flux.1 Schnell Full: Text to Image or Flux.1 Krea Dev. For the all-in-one files, Flux.1 Schnell FP8 or Flux.1 Dev fp8: Text to Image.

  4. Check the loaders

    Load Diffusion Model gets the model; on 24 GB or less, set its weight type to fp8_e4m3fn. DualCLIPLoader gets CLIP-L and T5 with type flux. Load VAE gets ae.safetensors. For a GGUF, swap the model loader for Unet Loader (GGUF).

  5. Keep CFG at 1

    Dev and Schnell have no negative prompt. The KSampler’s CFG stays at 1; the prompt strength is the guidance value, 3.5 by default.

  6. Write a prompt and run

    Plain descriptive sentences work well. T5 reads up to 512 tokens on Dev and 256 on Schnell, so long prompts are fine.

Settings that work.

Dev and Krea Dev

Steps
20 to 50
Guidance
3.5
CFG
1
Sampler
euler
Scheduler
simple
Size
1024 × 1024
Negative
none
T5 length
512 tokens

BFL’s model card uses 50 steps at guidance 3.5. Comfy’s templates use 20 steps and have no FluxGuidance node, so ComfyUI’s default of 3.5 applies. Add a FluxGuidance node to change it. The latent is EmptySD3LatentImage. Stay near 1 megapixel: beyond about 2 megapixels, images degrade.

Schnell

Steps
4
CFG
1
Sampler
euler
Scheduler
simple

Guidance and CFG are different things. Guidance is a value baked into Dev’s training that you pass in as conditioning. CFG is the sampler’s classifier-free guidance, and on Dev and Schnell it must stay at 1, or images turn blurry and burnt. For a working negative prompt, use a de-distilled fine-tune at a real CFG and follow its page for steps.

How fast.

GPUSetupTime
RTX 4090Dev, fp8_e4m3fn, 1024 px14 s per image[1]
RTX 4070 12 GBDev, default type2.5 s per step[1]
RTX 3060 12 GBDev, default type5 s per step[1]
RTX 4070 SuperQ4_0 GGUF1.9 s per step[2]
RTX 2080 8 GBQ4_0 GGUF, fp8 T53.2 s per step[2]
Mac Studio 96 GBQ4_K_S, 12 steps355 s[2]

User reports. The 3060 figure is with NVIDIA’s system memory fallback turned off. On the 4090, the default weight type took up to 10 minutes, because the model spilled out of graphics memory.

When it goes wrong.

mat1 and mat2 shapes cannot be multiplied (1x1280 and 768x3072)
The DualCLIPLoader type is set to sdxl. Set it to flux.
Blurry, washed-out or burnt images
CFG above 1, often with a negative prompt, on a distilled model. Set CFG to 1 and leave the negative empty.
Very slow on a 24 GB card
The full-precision model and the fp16 T5 don’t fit together. Set the weight type to fp8_e4m3fn in Load Diffusion Model, or use the fp8 T5.
Trying to convert Float8_e4m3fn to the MPS backend
An fp8 file on a Mac. Use the full-precision files or a GGUF.
A black image
The VAE decoded in float16. Run the VAE in bf16 or fp32.
clip missing: ['text_projection.weight']
A harmless warning. Nothing to fix.
Token indices sequence length is longer than the specified maximum sequence length
CLIP-L stops at 77 tokens and says so. T5 still reads the whole prompt, so ignore it.
A LoRA looks weaker with a GGUF model
In low-VRAM mode, part of the LoRA isn’t applied with GGUF models. Give ComfyUI more room or use an fp8 model.

My Flux images in Forge come out blurry. What am I doing wrong? The answer: set CFG to 1 and don’t use a negative prompt.

Hugging Face, FLUX.1-dev

Is Q4 good enough, or do I need Q8? Where does the quality actually start to break?

Hugging Face, city96 FLUX.1-dev GGUF

Questions.

What’s the difference between Flux.1 Dev and Schnell?

Dev is guidance-distilled and makes better images in 20 to 50 steps. Schnell is timestep-distilled for 1 to 4 steps, so it’s much faster, and it’s Apache 2.0. Both use the same encoders and VAE.

Can I use Flux.1 Dev commercially?

Schnell, yes: it’s Apache 2.0. Dev and Krea Dev are under the FLUX [dev] Non-Commercial License. The model card says outputs may be used commercially, but the licence text conflicts with that, so ask Black Forest Labs before running Dev for paid client work.

Why doesn’t the negative prompt work in Flux?

Dev and Schnell are distilled and run at CFG 1, where the negative prompt has no effect. Raising CFG makes images blurry. De-distilled fine-tunes restore a real CFG and a working negative prompt.

Should I use T5 fp16 or fp8?

The fp8 T5 (5.2 GB) saves memory with little loss. ComfyUI’s own examples recommend the fp16 one (9.8 GB) if you have more than 32 GB of system RAM. Avoid low-bit T5 GGUFs: they noticeably hurt prompt understanding.

Which Flux.1 GGUF should I pick?

Q8_0 (12.7 GB) for quality close to the original, Q6_K or Q5_K_S as the middle ground, and Q4_0 or Q4_K_S for an 8 GB card. Quality holds up to Q3 and breaks at Q2.

Does Flux.1 run on a Mac?

Yes, in ComfyUI on Apple Silicon, with the full-precision files or a GGUF. fp8 files don’t work on a Mac. Expect it to be slow: an M1 Max took about two minutes for a four-step Schnell image.

Do Flux.1 LoRAs work on Krea Dev or Flux.2?

On Krea Dev they load, since it has the same architecture, with mixed results. On Flux.2 they don’t load at all: it’s a new architecture.

Sources: FLUX.1 Dev model card, FLUX.1 Schnell model card, FLUX.1 Krea Dev model card, ComfyUI Flux examples, FLUX.1-dev speed thread [1], Flux GGUF speed thread [2], licence thread, macOS thread, ComfyUI issue #6969.

HEISS UI

Flux.1, ready when you are.

HEISS UI runs Flux.1 Dev, Schnell and Krea Dev on the ComfyUI you already have. Pick a file and it finds the parts that go with it.

  • Every part, found for you. The text encoders and VAE Flux.1 needs are listed with their sizes and fetched in one go.
  • Drop in the file and it runs. It knows the model from the file itself, even renamed, and picks settings that work.
  • Your LoRAs, sorted. The ones made for this model come first, with their trigger words.
  • Upscale and compare. One click to upscale with SeedVR2, then drag a slider to see what changed.

Kontext isn’t built in; it runs as your own workflow. Free and open source. macOS, Windows and Linux. Runs on your ComfyUI.

HEISS UI with a gallery of generated images and the prompt composer at the bottom.